proxy_spearman__PXR__T0.50__C0.25__INTERNAL — voice-corrected

Manifest tier. proxy_spearman, rule PXR, T=0.5, step cap 0.25. Population 1,016 chains (11 h) over 6 corpora. The SHAREABLE variant of this tier (podcast and evasnippets excluded) holds 901.

This is the voice-conversion-corrected version of this tier. The original, uncorrected page is still there and unchanged: t_proxy_spearman__PXR__T0.50__C0.25__INTERNAL.html. Every card below carries both renders so you can switch between them without leaving the page. What was done, and what it measured →
20chains converted
74segments re-voiced
0.776 → 0.770median worst-to-anchor identity cosine
93 %median emotion delta retained
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken.

(underlined, plain · delivery, style) — the tag before the words. Emotions first, then how it is delivered. Underlined descriptors are the ones that change across this chain — anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words are a real non-speech sound, printed where it happens.

The full generated caption for any clip is under “full caption & clip details”. Its perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.

The tags and captions describe the ORIGINAL clips — they are the corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the converted audio are the ones printed in each card's own paragraph and chips.
Pain ↓  /  Intoxication Altered States of Consciousnessidentity −0.06 emotion 50 %   proxy_spearman__PXR__T0.50__C0.25__INTERNAL · #1

This chain comes from the proxy rule: the same two-sided test as above, but because Intoxication Altered States of Consciousness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Intoxication Altered States of Consciousness barely there — 0.22, lower than 78 % of clips in this corpus — and ends with it clearly present at 0.73, higher than 73 % of clips in this corpus. That is a total rise of 0.50.

At the same time Pain goes the other way, from 0.90 (higher than 90 % of clips in this corpus) to 0.35 (lower than 65 % of clips in this corpus), a change of -0.55. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are -0.04, then +0.23, then +0.17, then +0.14 — not a clean run: step 1 moves back the other way by 0.04 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.85 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.87 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.85. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 36 s · ko · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.829 before conversion and 0.769 after — it fell by 0.060. Neighbour-to-neighbour the worst pair went 0.720 → 0.705. (The earlier render, with segment 1 left raw, scores 0.699 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Intoxication Altered States of Consciousness moved +0.504 in the original and +0.254 after conversion — 50 % of the delta retained. On the other named axis, Pain, -0.549 became +0.196.

Quality. Mean predicted overall quality across the segments went 3.00 → 3.13 (+0.13) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.829 → 0.769 -0.060identity cos neighbours 0.720 → 0.705d_b rescored +0.504 → +0.254d_a rescored -0.549 → +0.196d_a mined -0.549d_b mined 0.504min_cos_consec (site) 0.8660min_cos_anchor (site) 0.8469dataset emolialang kospeaker KO_zzLpjKEJ-D8total 34.6schain gain +2.3 dBseam step 1.2 dBcrossfades 100/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed
(normal-paced, fairly steady, little disfluency, authoritative) 마릴린 몰로가 사망하게 된 원인은 아직까지 명확히 밝혀진 것은 없습니다만,
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: authoritative, formal; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 0.0/10; 6.6s, KO.
KO_zzLpjKEJ-D8_W000013 · in -17.3 dBFS · gain -2.7 dB · emolia-03140
(measured, steady, no disfluency, formal) 극심한 우울증으로 인한 과다, 약물 복용이 원인이 아니었을까 하는 추측이 가장 많습니다.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.9/10; 7.6s, KO.
KO_zzLpjKEJ-D8_W000014 · in -18.7 dBFS · gain -1.3 dB · emolia-03140
(normal-paced, fairly steady, no disfluency, formal) 마릴린 물론은, 알고 계시는 것처럼 모든 남성들의 사랑을 받았습니다.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 0.0/10; 5.3s, KO.
KO_zzLpjKEJ-D8_W000015 · in -16.8 dBFS · gain -3.2 dB · emolia-03140
(thankfulness gratitude, relief · measured, steady, no disfluency, monologue) 특히, 어느 셀러리맨의 죽음이라는 작품으로 퓨리처상을 받은 그 당시의 스타작가 아스 밀러가 가장 사랑했던 여자였기도 했습니다.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as thankfulness gratitude, relief; style: monologue, narration; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.6/10; 12.6s, KO.
KO_zzLpjKEJ-D8_W000016 · in -18.8 dBFS · gain -1.2 dB · emolia-03140
(normal-paced, fairly steady, no disfluency, formal) 실제로 두 사람은 아주 뜨겁게 사랑을 했고,
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, didactic; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 1.3/10; 3.3s, KO.
KO_zzLpjKEJ-D8_W000017 · in -17.5 dBFS · gain -2.5 dB · emolia-03140
Anger ↓  /  Concentrationidentity +0.06 emotion 114 %   proxy_spearman__PXR__T0.50__C0.25__INTERNAL · #2

This chain comes from the proxy rule: the same two-sided test as above, but because Concentration is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Concentration below average — 0.37, lower than 63 % of clips in this corpus — and ends with it strongly present at 0.88, higher than 88 % of clips in this corpus. That is a total rise of 0.51.

At the same time Anger goes the other way, from 0.88 (higher than 88 % of clips in this corpus) to 0.28 (lower than 72 % of clips in this corpus), a change of -0.61. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.19, then +0.20, then +0.20, then -0.08 — not a clean run: step 4 moves back the other way by 0.08 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.79 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.84 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.79, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 48 s · zh · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.622 before conversion and 0.685 after — it rose by 0.063. Neighbour-to-neighbour the worst pair went 0.840 → 0.768. (The earlier render, with segment 1 left raw, scores 0.590 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.509 in the original and +0.581 after conversion — 114 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Anger, -0.606 became -0.591.

Quality. Mean predicted overall quality across the segments went 3.14 → 3.18 (+0.04) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.622 → 0.685 +0.063identity cos neighbours 0.840 → 0.768d_b rescored +0.509 → +0.581d_a rescored -0.606 → -0.591d_a mined -0.606d_b mined 0.511min_cos_consec (site) 0.8414min_cos_anchor (site) 0.7898dataset emolialang zhspeaker ZH_B00058_S06018total 46.6schain gain +1.9 dBseam step 1.5 dBcrossfades 150/100/100/150 ms
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, no background noise, normally alert, slightly relaxed, fairly steady, moderate pitch range
(normal-paced, some disfluency, average clarity, casual) 他到了蜀蜀汉之后呢,到了这个四川这个地方。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, conversational; good recording, no background noise; genuineness 3.1/6; vocal-burst blend 4.0/10; 3.3s, ZH.
ZH_B00058_S06018_W000038 · in -22.5 dBFS · gain +2.5 dB · emolia-03860
(measured, no disfluency, clear, narration) 一个北方人到达南方人,最大的困境就是如何把南方人的心收过来。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, formal; average recording, no background noise; genuineness 1.6/6; vocal-burst blend 3.1/10; 6.0s, ZH.
ZH_B00058_S06018_W000039 · in -21.5 dBFS · gain +1.5 dB · emolia-03860
(measured, some disfluency, clear, monologue) 所以孔明做了很多的动作,设计了很多的制度,做了很多的机制,也一步一步的把南方人思想团结在以刘备为中心的政府身边。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, narration; average recording, no background noise; genuineness 2.2/6; vocal-burst blend 5.1/10; 11.2s, ZH.
ZH_B00058_S06018_W000040 · in -21.1 dBFS · gain +1.1 dB · emolia-03860
(contemplation, concentration, pride · measured, some disfluency, clear, monologue) (ahem) (ahem) (ahem) 那这种做法我们称为工心为上。其实公司呢在做事的过程之中,仅仅发工资不见得把人心留住啊,必须要通过一套机制呢,把真正的啊这种心呢啊给他收住,而心呢是看不见摸不到的,对方怎么想也不清楚。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation, concentration, pride; style: monologue, didactic; average recording, no background noise; genuineness 1.8/6; vocal-burst blend 3.8/10; 17.0s, ZH.
ZH_B00058_S06018_W000041 · in -20.7 dBFS · gain +0.7 dB · emolia-03860
(measured, no disfluency, clear, monologue) 所以呢我我希望大家呢千万要记住一句话叫身在曹营心在汉,尤其是刚过完春节,很多人呢可能打算呢离开公司。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, narration; average recording, no background noise; genuineness 1.3/6; vocal-burst blend 3.8/10; 9.8s, ZH.
ZH_B00058_S06018_W000042 · in -20.7 dBFS · gain +0.7 dB · emolia-03860
Thankfulness Gratitude ↓  /  Concentrationidentity −0.14 emotion 133 %   proxy_spearman__PXR__T0.50__C0.25__INTERNAL · #3

This chain comes from the proxy rule: the same two-sided test as above, but because Concentration is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Concentration below average — 0.40, lower than 60 % of clips in this corpus — and ends with it at the very top of the corpus at 0.92, higher than 92 % of clips in this corpus. That is a total rise of 0.51.

At the same time Thankfulness Gratitude goes the other way, from 0.73 (higher than 73 % of clips in this corpus) to 0.19 (lower than 81 % of clips in this corpus), a change of -0.54. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.23, then +0.15, then +0.07, then +0.06 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.66 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.66 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.66, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 33 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.669 before conversion and 0.527 after — it fell by 0.142. Neighbour-to-neighbour the worst pair went 0.669 → 0.527. (The earlier render, with segment 1 left raw, scores 0.467 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.510 in the original and +0.679 after conversion — 133 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Thankfulness Gratitude, -0.543 became -0.377.

Quality. Mean predicted overall quality across the segments went 2.68 → 2.95 (+0.27) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.669 → 0.527 -0.142identity cos neighbours 0.669 → 0.527d_b rescored +0.510 → +0.679d_a rescored -0.543 → -0.377d_a mined -0.543d_b mined 0.510min_cos_consec (site) 0.6637min_cos_anchor (site) 0.6637dataset emolialang enspeaker EN_Z0y2GAYzZbktotal 31.5schain gain -0.9 dBseam step 4.8 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 4 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, slightly dark, quiet background, measured, frequent disfluency, fairly narrow pitch
(normally alert, slightly relaxed, fairly steady, casual) To (ahem) (ahem) enable the appropriate coordination, okay.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, thin; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, conversational; below-average recording, quiet background; genuineness 5.2/6; vocal-burst blend 0.4/10; 3.3s, EN.
EN_Z0y2GAYzZbk_W000010 · in -11.5 dBFS · gain -8.5 dB · emolia-02555
(normally alert, slightly relaxed, steady, didactic) So we started with the data regarding the circuit breakers.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; slurred, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: didactic, monologue; average recording, quiet background; genuineness 3.3/6; vocal-burst blend 0.8/10; 5.5s, EN.
EN_Z0y2GAYzZbk_W000011 · in -9.1 dBFS · gain -10.9 dB · emolia-02555
(normally alert, slightly relaxed, steady, monologue) So if you look at the circuit breakers, (low mumble) (low mumble) we had mentioned that the circuit breakers
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, didactic; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 0.4/10; 6.9s, EN.
EN_Z0y2GAYzZbk_W000012 · in -9.1 dBFS · gain -10.9 dB · emolia-02555
(fatigue exhaustion, emotional numbness · subdued, slightly relaxed, steady, monologue) Definite time delay of (resonant hum) uh, 40 milliseconds or two cycles, (low mumble) uh, it has extremely inverse
full caption & clip details
A young adult masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as fatigue exhaustion, emotional numbness; style: monologue, didactic; average recording, quiet background; genuineness 3.3/6; vocal-burst blend 0.3/10; 8.1s, EN.
EN_Z0y2GAYzZbk_W000013 · in -11.2 dBFS · gain -8.8 dB · emolia-02555
(concentration · very low-energy, relaxed, fairly steady, monologue) Of 30 seconds and we (low mumble) mention that the I square t (low mumble) for conductor
full caption & clip details
A young adult masculine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, slightly thin; slurred, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, neutral openness; reads as concentration; style: monologue, didactic; below-average recording, quiet background; genuineness 3.8/6; vocal-burst blend 0.4/10; 8.6s, EN.
EN_Z0y2GAYzZbk_W000015 · in -9.9 dBFS · gain -10.1 dB · emolia-02555
Thankfulness Gratitude ↓  /  Painidentity +0.04 emotion 99 %   proxy_spearman__PXR__T0.50__C0.25__INTERNAL · #4

This chain comes from the proxy rule: the same two-sided test as above, but because Pain is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Pain barely there — 0.20, lower than 80 % of clips in this corpus — and ends with it at the very top of the corpus at 0.91, higher than 91 % of clips in this corpus. That is a total rise of 0.71.

At the same time Thankfulness Gratitude goes the other way, from 0.70 (higher than 70 % of clips in this corpus) to 0.17 (lower than 83 % of clips in this corpus), a change of -0.53. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.15, then +0.20, then +0.17, then +0.19 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.83 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.81 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.83. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 18 s · zh · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.829 before conversion and 0.868 after — it rose by 0.039. Neighbour-to-neighbour the worst pair went 0.809 → 0.842. (The earlier render, with segment 1 left raw, scores 0.830 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Pain moved +0.714 in the original and +0.708 after conversion — 99 % of the delta retained, which is essentially all of it. On the other named axis, Thankfulness Gratitude, -0.535 became -0.362.

Quality. Mean predicted overall quality across the segments went 2.84 → 2.83 (-0.01) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.829 → 0.868 +0.039identity cos neighbours 0.809 → 0.842d_b rescored +0.714 → +0.708d_a rescored -0.535 → -0.362d_a mined -0.535d_b mined 0.714min_cos_consec (site) 0.8062min_cos_anchor (site) 0.8313dataset emolialang zhspeaker ZH_B00080_S09132total 16.7schain gain +0.3 dBseam step 0.9 dBcrossfades 100/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, no background noise, normally alert, slightly relaxed, fairly steady
(measured, narration, monologue) 为了吸引过客,村子四周都挂满了招牌。
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, monologue; average recording, no background noise; genuineness 1.9/6; vocal-burst blend 2.5/10; 3.2s, ZH.
ZH_B00080_S09132_W000014 · in -22.9 dBFS · gain +2.9 dB · emolia-04074
(measured, formal, monologue) 我甚至常常闯入别人家里,还受到了很好的招待呢。
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 3.9/10; 3.5s, ZH.
ZH_B00080_S09132_W000015 · in -25.7 dBFS · gain +5.7 dB · emolia-04074
(measured, monologue, storytelling) 当我在村子里待到很晚的时候,要在黑夜中赶路。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, storytelling; good recording, no background noise; genuineness 1.7/6; vocal-burst blend 4.6/10; 3.6s, ZH.
ZH_B00080_S09132_W000016 · in -25.1 dBFS · gain +5.1 dB · emolia-04074
(measured, storytelling, narration) 那是件十分过瘾的事,尤其在夜黑风高的晚上。
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: storytelling, narration; average recording, no background noise; genuineness 2.5/6; vocal-burst blend 5.0/10; 3.6s, ZH.
ZH_B00080_S09132_W000017 · in -24.8 dBFS · gain +4.8 dB · emolia-04074
(pain · normal-paced, monologue, formal) 我从村中某个明亮的客厅或演讲厅起航。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pain; style: monologue, formal; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 4.6/10; 3.4s, ZH.
ZH_B00080_S09132_W000018 · in -24.5 dBFS · gain +4.5 dB · emolia-04074
Concentration ↓  /  Intoxication Altered States of Consciousnessidentity −0.10 emotion 89 %   proxy_spearman__PXR__T0.50__C0.25__INTERNAL · #5

This chain comes from the proxy rule: the same two-sided test as above, but because Intoxication Altered States of Consciousness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Intoxication Altered States of Consciousness below average — 0.31, lower than 69 % of clips in this corpus — and ends with it strongly present at 0.82, higher than 82 % of clips in this corpus. That is a total rise of 0.51.

At the same time Concentration goes the other way, from 0.84 (higher than 84 % of clips in this corpus) to 0.17 (lower than 83 % of clips in this corpus), a change of -0.67. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.22, then +0.22, then +0.04, then +0.03 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.78 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.71 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.78, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 28 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.702 before conversion and 0.597 after — it fell by 0.104. Neighbour-to-neighbour the worst pair went 0.752 → 0.654. (The earlier render, with segment 1 left raw, scores 0.580 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Intoxication Altered States of Consciousness moved +0.507 in the original and +0.452 after conversion — 89 % of the delta retained, which is most of it. On the other named axis, Concentration, -0.667 became -0.685.

Quality. Mean predicted overall quality across the segments went 2.79 → 2.96 (+0.17) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.702 → 0.597 -0.104identity cos neighbours 0.752 → 0.654d_b rescored +0.507 → +0.452d_a rescored -0.667 → -0.685d_a mined -0.667d_b mined 0.507min_cos_consec (site) 0.7131min_cos_anchor (site) 0.7777dataset emolialang enspeaker EN_DuD2CAiU3lEtotal 26.5schain gain +1.4 dBseam step 1.5 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed, clear
(normal-paced, fairly steady, no disfluency, formal) The person icon here relates to groups, again something we don't focus on in this MOOC but which you can read about in the documentation.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.3/10; 8.0s, EN.
EN_DuD2CAiU3lE_W000023 · in -19.6 dBFS · gain -0.4 dB · emolia-02536
(relief · normal-paced, fairly steady, little disfluency, formal) When we click to hide it, we can see now that the eye icon has a line through it and the announcements are greyed out.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as relief; style: formal, whispered; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.2/10; 7.0s, EN.
EN_DuD2CAiU3lE_W000024 · in -16.4 dBFS · gain -3.6 dB · emolia-02536
(normal-paced, fairly steady, no disfluency, storytelling) If we go to our user menu, we see Switch Roll To.
full caption & clip details
A child feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, no disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; no dominant emotion; style: storytelling, formal; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 1.9/10; 3.8s, EN.
EN_DuD2CAiU3lE_W000025 · in -15.5 dBFS · gain -4.5 dB · emolia-02536
(normal-paced, steady, no disfluency, formal) And as a teacher, we can switch our role amongst others to a student.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.7/10; 4.8s, EN.
EN_DuD2CAiU3lE_W000026 · in -18.8 dBFS · gain -1.2 dB · emolia-02536
(measured, fairly steady, no disfluency, storytelling) Or, as we've renamed it in our course, Learner.
full caption & clip details
A child feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; style: storytelling, formal; good recording, no background noise; no dominant emotion; genuineness 1.0/6; vocal-burst blend 2.7/10; 3.7s, EN.
EN_DuD2CAiU3lE_W000027 · in -17.0 dBFS · gain -3.0 dB · emolia-02536
Thankfulness Gratitude ↓  /  Infatuationidentity −0.11 emotion 64 %   proxy_spearman__PXR__T0.50__C0.25__INTERNAL · #6

This chain comes from the proxy rule: the same two-sided test as above, but because Infatuation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Infatuation below average — 0.27, lower than 73 % of clips in this corpus — and ends with it strongly present at 0.80, higher than 80 % of clips in this corpus. That is a total rise of 0.52.

At the same time Thankfulness Gratitude goes the other way, from 0.66 (higher than 66 % of clips in this corpus) to 0.15 (lower than 85 % of clips in this corpus), a change of -0.51. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.14, then +0.18, then +0.00, then +0.20 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.92 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.94 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.92), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 45 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.894 before conversion and 0.785 after — it fell by 0.110. Neighbour-to-neighbour the worst pair went 0.851 → 0.725. (The earlier render, with segment 1 left raw, scores 0.565 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Infatuation moved +0.524 in the original and +0.335 after conversion — 64 % of the delta retained. On the other named axis, Thankfulness Gratitude, -0.510 became -0.519.

Quality. Mean predicted overall quality across the segments went 2.91 → 3.03 (+0.12) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.894 → 0.785 -0.110identity cos neighbours 0.851 → 0.725d_b rescored +0.524 → +0.335d_a rescored -0.510 → -0.519d_a mined -0.510d_b mined 0.524min_cos_consec (site) 0.9386min_cos_anchor (site) 0.9226dataset emolialang enspeaker EN_92xztpjhoJ4total 43.5schain gain +1.1 dBseam step 1.0 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(steady, formal, authoritative) The Chicago Tribune is a daily newspaper based in Chicago, Illinois, United States, owned by Tribune publishing
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; mildly explicit content; genuineness 0.2/6; vocal-burst blend 0.0/10; 6.8s, EN.
EN_92xztpjhoJ4_W000000 · in -14.2 dBFS · gain -5.8 dB · emolia-02017
(fairly steady, newsreading, formal) Founded in 1847, and formerly self-styled as the ''world's greatest newspaper'' for which WGN Radio and Television are named, it remains the most read daily newspaper of the Chicago metropolitan area and the Great Lakes region.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 13.4s, EN.
EN_92xztpjhoJ4_W000001 · in -14.2 dBFS · gain -5.8 dB · emolia-02017
(fairly steady, formal, newsreading) The Tribune announced it would continue publishing as a broadsheet for home delivery, but would publish in tabloid format for news stand, news box, and commuter station sales
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 8.9s, EN.
EN_92xztpjhoJ4_W000003 · in -15.5 dBFS · gain -4.5 dB · emolia-02017
(fairly steady, formal, newsreading) == History == === Beginnings === The Tribune was founded by James Kelly, John E. Wheeler, and Joseph K. C. Forrest, publishing the first edition on June 10, 1847
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 10.4s, EN.
EN_92xztpjhoJ4_W000005 · in -14.6 dBFS · gain -5.4 dB · emolia-02017
(fairly steady, formal, authoritative) Numerous changes in ownership and editorship took place over the next eight years.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 0.5/10; 4.7s, EN.
EN_92xztpjhoJ4_W000006 · in -14.3 dBFS · gain -5.7 dB · emolia-02017
Elation ↓  /  Contemplationidentity +0.43 emotion 95 %   proxy_spearman__PXR__T0.50__C0.25__INTERNAL · #7

This chain comes from the proxy rule: the same two-sided test as above, but because Contemplation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Contemplation below average — 0.40, lower than 60 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.54.

At the same time Elation goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.46 (lower than 54 % of clips in this corpus), a change of -0.53. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.14, then +0.19, then +0.22 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.02 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.04 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.02, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 71 s · en · podcast

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.003 before conversion and 0.437 after — it rose by 0.434. Neighbour-to-neighbour the worst pair went 0.012 → 0.506. (The earlier render, with segment 1 left raw, scores 0.460 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contemplation moved +0.543 in the original and +0.517 after conversion — 95 % of the delta retained, which is essentially all of it. On the other named axis, Elation, -0.862 became -0.859.

Quality. Mean predicted overall quality across the segments went 3.00 → 3.16 (+0.16) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.003 → 0.437 +0.434identity cos neighbours 0.012 → 0.506d_b rescored +0.543 → +0.517d_a rescored -0.862 → -0.859d_a mined -0.528d_b mined 0.542min_cos_consec (site) 0.0388min_cos_anchor (site) 0.0242dataset podcastlang enspeaker 946087total 69.8schain gain +5.5 dBseam step 1.6 dBcrossfades 150/150/100 ms
Script — 4 chunks, 3 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · neutral-toned, fairly smooth, balanced body, normally alert, some disfluency, light breath
(elation, pleasure ecstasy, infatuation · brisk, neutral tension, moderately variable, casual) actually, yeah, I did get that feedback from someone (low mumble) uh that (ahem) uh he he was just auditioning for a project for fun in one of my Discord servers. And I auditioned for like, yeah, the the regular protagonist, like, actually, I think you could be like her like her snarky best friend, which is like a little a little bit pushy, but she's also like deeply nice. And I actually ended up recording for that character, and he said I did really well. I'm really excited to hear how that came out.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as elation, pleasure ecstasy, infatuation; style: casual, conversational; good recording, quiet background; genuineness 3.8/6; vocal-burst blend 6.5/10; 24.7s, EN.
946087_00272360 · in -28.6 dBFS · gain +8.6 dB · podcast-03290
(pride, embarrassment, triumph · normal-paced, slightly relaxed, fairly steady, casual) Walter said when I started doing hair, I would do haircuts for five dollars because otherwise I would make zero. Now I get 40 for a haircut, and over time as I get busy, I raise my prices. It's about increasing quality of clients, not number of clients. Uh
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as pride, embarrassment, triumph; style: casual, monologue; good recording, no background noise; genuineness 1.8/6; vocal-burst blend 1.6/10; 13.2s, EN.
946087_00275064 · in -34.2 dBFS · gain +14.2 dB · podcast-06213
(disappointment, impatience and irritability, anger · normal-paced, neutral tension, moderately variable, casual) you can't get there unless you're working. So nothing wrong with doing it for less. If it's the alternative, is doing nothing at all. (surprised gasp) I mean it's it's it's that catch twenty-two that we're all in. We we can't raise our prices without experience, we can't get experience without gigs, we can't get gigs because everybody undercover
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as disappointment, impatience and irritability, anger; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.3/6; vocal-burst blend 3.1/10; 19.2s, EN.
946087_00276576 · in -34.0 dBFS · gain +14.0 dB · podcast-03279
(contemplation · normal-paced, slightly relaxed, fairly steady, casual) well when you when you start seeing it from a business perspective, it it does make sense, you know. Like businesses that don't have a lot of (ahem) uh clientele to work with are going to offer cheaper rates, and
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation; style: casual, conversational; good recording, no background noise; genuineness 4.1/6; vocal-burst blend 3.3/10; 13.2s, EN.
946087_00278744 · in -29.6 dBFS · gain +9.6 dB · podcast-03276
Malevolence Malice ↓  /  Thankfulness Gratitudeidentity +0.01 emotion 103 %   proxy_spearman__PXR__T0.50__C0.25__INTERNAL · #8

This chain comes from the proxy rule: the same two-sided test as above, but because Thankfulness Gratitude is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Thankfulness Gratitude barely there — 0.21, lower than 79 % of clips in this corpus — and ends with it strongly present at 0.76, higher than 76 % of clips in this corpus. That is a total rise of 0.54.

At the same time Malevolence Malice goes the other way, from 0.94 (higher than 94 % of clips in this corpus) to 0.44 (lower than 56 % of clips in this corpus), a change of -0.51. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.24, then -0.06, then +0.12, then +0.24 — not a clean run: step 2 moves back the other way by 0.06 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.95 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.94 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.95), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 41 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.803 before conversion and 0.817 after — it rose by 0.014. Neighbour-to-neighbour the worst pair went 0.898 → 0.829. (The earlier render, with segment 1 left raw, scores 0.644 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Thankfulness Gratitude moved +0.566 in the original and +0.583 after conversion — 103 % of the delta retained, which is essentially all of it. On the other named axis, Malevolence Malice, -0.509 became -0.460.

Quality. Mean predicted overall quality across the segments went 2.99 → 3.10 (+0.11) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.803 → 0.817 +0.014identity cos neighbours 0.898 → 0.829d_b rescored +0.566 → +0.583d_a rescored -0.509 → -0.460d_a mined -0.509d_b mined 0.542min_cos_consec (site) 0.9434min_cos_anchor (site) 0.9492dataset emolialang enspeaker EN_B00008_S04346total 39.7schain gain +3.9 dBseam step 1.3 dBcrossfades 150/100/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(malevolence malice · steady, formal, newsreading) After Sun Quan limped back to camp under the protection of his officers, he handsomely rewarded Ling Tong for his bravery, as well as the lieutenant who told him to jump across the bridge.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as malevolence malice; style: formal, newsreading; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.0/10; 11.3s, EN.
EN_B00008_S04346_W000086 · in -21.0 dBFS · gain +1.0 dB · emolia-00419
(fairly steady, formal, didactic) Sun Quan then had his army fall back to Ruxu, where they regrouped, sent word back to the Southlands for reinforcements, and began planning another attack by land and by water.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, didactic; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.1/10; 11.6s, EN.
EN_B00008_S04346_W000087 · in -19.2 dBFS · gain -0.8 dB · emolia-00419
(fairly steady, formal, narration) Inside Hefei, the victorious Zhang Liao soon learned that Sun Quan and his army were not late yet.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, narration; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 0.7/10; 6.4s, EN.
EN_B00008_S04346_W000088 · in -20.2 dBFS · gain +0.2 dB · emolia-00419
(distress, fear, shame · fairly steady, formal, monologue) And he was worried about not having enough troops to withstand another siege.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as distress, fear, shame; style: formal, monologue; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 1.2/10; 4.4s, EN.
EN_B00008_S04346_W000089 · in -22.6 dBFS · gain +2.6 dB · emolia-00419
(fairly steady, formal, didactic) So Zhang Liao quickly sent word to the region of Hanzhong, where Cao Cao was currently located, to ask for help.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, didactic; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.7/10; 6.7s, EN.
EN_B00008_S04346_W000090 · in -21.4 dBFS · gain +1.4 dB · emolia-00419
Hope Enthusiasm Optimism ↓  /  Distressidentity +0.05 emotion 100 %   proxy_spearman__PXR__T0.50__C0.25__INTERNAL · #9

This chain comes from the proxy rule: the same two-sided test as above, but because Distress is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Distress around average — 0.44, lower than 56 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.55.

At the same time Hope Enthusiasm Optimism goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.45 (lower than 55 % of clips in this corpus), a change of -0.53. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.00, then +0.00, then +0.46, then +0.10 — a plateau around step 1, where it barely moves.

The largest step is 0.46, which is above the 0.25 cap the strict rule would impose — worth knowing when judging how gradual it sounds.

Same speaker? The least similar clip scores 0.73 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.69 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.73, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 43 s · en · podcast

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.748 before conversion and 0.797 after — it rose by 0.049. Neighbour-to-neighbour the worst pair went 0.711 → 0.746. (The earlier render, with segment 1 left raw, scores 0.710 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Distress moved +0.554 in the original and +0.554 after conversion — 100 % of the delta retained, which is essentially all of it. On the other named axis, Hope Enthusiasm Optimism, -0.526 became -0.708.

Quality. Mean predicted overall quality across the segments went 2.80 → 2.99 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.748 → 0.797 +0.049identity cos neighbours 0.711 → 0.746d_b rescored +0.554 → +0.554d_a rescored -0.526 → -0.708d_a mined -0.526d_b mined 0.554min_cos_consec (site) 0.6931min_cos_anchor (site) 0.7304dataset podcastlang enspeaker 70041total 41.6schain gain +1.6 dBseam step 6.2 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, fairly smooth, normally alert, some disfluency, average clarity, light breath
(hope enthusiasm optimism · normal-paced, slightly relaxed, moderately variable, casual) online that you can get astrology I feel like is a little bit more accessible in that
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as hope enthusiasm optimism; style: casual, conversational; good recording, no background noise; genuineness 4.2/6; vocal-burst blend 4.2/10; 5.2s, EN.
70041_00234544 · in -21.3 dBFS · gain +1.3 dB · podcast-05577
(thankfulness gratitude · brisk, slightly relaxed, fairly steady, conversational) there's great resources like you can go on gene key's website he has a lot of information there so and the book is
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as thankfulness gratitude; style: conversational, casual; good recording, no background noise; genuineness 3.3/6; vocal-burst blend 5.8/10; 7.0s, EN.
70041_00235352 · in -21.4 dBFS · gain +1.4 dB · podcast-05575
(elation · brisk, neutral tension, moderately variable, conversational) wonderful as well so there's (ahem) a definitely a lot of resources for parents if they wanted to and always they can reach out to you which I had good
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as elation; style: conversational, casual; good recording, quiet background; genuineness 4.5/6; vocal-burst blend 5.5/10; 8.6s, EN.
70041_00236048 · in -20.9 dBFS · gain +0.9 dB · podcast-05571
(awe, confusion, astonishment surprise · normal-paced, slightly relaxed, fairly steady, casual) In the (low mumble) zodiac.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as awe, confusion, astonishment surprise; style: casual, conversational; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 3.7/10; 11.9s, EN.
70041_00269466 · in -20.8 dBFS · gain +0.8 dB · podcast-03390
(distress, fear, pain · normal-paced, slightly relaxed, fairly steady, casual) These are very deep wounds about our identity that needs to come to the surface. For some, this can look like an ego death, a
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as distress, fear, pain; style: casual, monologue; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 2.8/10; 9.6s, EN.
70041_00270704 · in -21.3 dBFS · gain +1.3 dB · podcast-03397
Triumph ↓  /  Concentrationidentity −0.04 emotion 116 %   proxy_spearman__PXR__T0.50__C0.25__INTERNAL · #10

This chain comes from the proxy rule: the same two-sided test as above, but because Concentration is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Concentration below average — 0.25, lower than 75 % of clips in this corpus — and ends with it strongly present at 0.79, higher than 79 % of clips in this corpus. That is a total rise of 0.54.

At the same time Triumph goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.38 (lower than 62 % of clips in this corpus), a change of -0.61. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.18, then +0.24, then +0.12 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.96 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.96 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.96), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 39 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.880 before conversion and 0.845 after — it fell by 0.035. Neighbour-to-neighbour the worst pair went 0.813 → 0.800. (The earlier render, with segment 1 left raw, scores 0.586 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.539 in the original and +0.626 after conversion — 116 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Triumph, -0.613 became -0.549.

Quality. Mean predicted overall quality across the segments went 2.88 → 3.05 (+0.17) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.880 → 0.845 -0.035identity cos neighbours 0.813 → 0.800d_b rescored +0.539 → +0.626d_a rescored -0.613 → -0.549d_a mined -0.613d_b mined 0.539min_cos_consec (site) 0.9604min_cos_anchor (site) 0.9604dataset emolialang enspeaker EN_p-GGhAlaWtEtotal 37.5schain gain +0.8 dBseam step 2.9 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, good recording, no background noise, normal-paced, normally alert, slightly relaxed
(triumph, pride · fairly steady, almost no disfluency, formal, narration) My class was the first class to have women in it, it was the first class to have a significant effort to recruit African Americans
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as triumph, pride; style: formal, narration; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 7.8s, EN.
EN_p-GGhAlaWtE_W000156 · in -14.7 dBFS · gain -5.3 dB · emolia-01803
(awe, astonishment surprise, contemplation · steady, no disfluency, formal, narration) It was an extraordinary time, and in that span of time is the change of an entire generation
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as awe, astonishment surprise, contemplation; style: formal, narration; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.3/10; 5.5s, EN.
EN_p-GGhAlaWtE_W000157 · in -13.2 dBFS · gain -6.8 dB · emolia-01803
(fairly steady, no disfluency, newsreading, formal) In 2009, former British Prime Minister Tony Blair picked Yale as one location, the others are Britain's Durham University and University Technologie Mara, for the Tony Blair Faith Foundation's United States Faith and Globalization Initiative
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 13.9s, EN.
EN_p-GGhAlaWtE_W000158 · in -13.5 dBFS · gain -6.5 dB · emolia-01803
(steady, no disfluency, newsreading, formal) As of 2009, former Mexican President Ernesto Zedillo is the director of the Yale Center for the Study of Globalization and teaches an undergraduate seminar,
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: newsreading, formal; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.0/10; 10.9s, EN.
EN_p-GGhAlaWtE_W000159 · in -14.6 dBFS · gain -5.4 dB · emolia-01803
Intoxication Altered States of Consciousness ↓  /  Emotional Numbnessidentity −0.10 emotion 76 %   proxy_spearman__PXR__T0.50__C0.25__INTERNAL · #11

This chain comes from the proxy rule: the same two-sided test as above, but because Emotional Numbness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Emotional Numbness barely there — 0.13, lower than 87 % of clips in this corpus — and ends with it strongly present at 0.83, higher than 83 % of clips in this corpus. That is a total rise of 0.71.

At the same time Intoxication Altered States of Consciousness goes the other way, from 0.69 (higher than 69 % of clips in this corpus) to 0.08 (lower than 92 % of clips in this corpus), a change of -0.61. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.10, then +0.22, then +0.14, then +0.24 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.87 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.87 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.87. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 41 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.874 before conversion and 0.771 after — it fell by 0.103. Neighbour-to-neighbour the worst pair went 0.883 → 0.771. (The earlier render, with segment 1 left raw, scores 0.536 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.705 in the original and +0.536 after conversion — 76 % of the delta retained, which is most of it. On the other named axis, Intoxication Altered States of Consciousness, -0.606 became -0.483.

Quality. Mean predicted overall quality across the segments went 2.82 → 2.93 (+0.11) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.874 → 0.771 -0.103identity cos neighbours 0.883 → 0.771d_b rescored +0.705 → +0.536d_a rescored -0.606 → -0.483d_a mined -0.606d_b mined 0.705min_cos_consec (site) 0.8652min_cos_anchor (site) 0.8652dataset emolialang enspeaker EN_SKjE_MI17TAtotal 39.2schain gain +3.5 dBseam step 1.8 dBcrossfades 150/150/100/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed
(normal-paced, fairly steady, moderate pitch range, formal) Jam sandwiches are thought to have originated at around the 19th century in the United Kingdom.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 1.0/10; 5.2s, EN.
EN_SKjE_MI17TA_W000001 · in -15.0 dBFS · gain -5.0 dB · emolia-01762
(normal-paced, steady, moderate pitch range, formal) In Scotland, they are also known as pieces and jam, or geely pieces.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 1.1/10; 4.1s, EN.
EN_SKjE_MI17TA_W000002 · in -15.5 dBFS · gain -4.5 dB · emolia-01762
(normal-paced, fairly steady, moderate pitch range, formal) The jam sandwich was an affordable food which was a major part of the diets of the lower, working class people of cities such as London and Glasgow.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, newsreading; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.0/10; 7.9s, EN.
EN_SKjE_MI17TA_W000003 · in -15.2 dBFS · gain -4.8 dB · emolia-01762
(normal-paced, fairly steady, moderate pitch range, newsreading) One plausible reason for this was that the ingredients that the jam sandwiches were made from cost little to manufacture and due to taxes being lifted on sugar in 1880, it became widely available as a cheap foodstuff.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 12.3s, EN.
EN_SKjE_MI17TA_W000004 · in -15.3 dBFS · gain -4.7 dB · emolia-01762
(measured, steady, fairly narrow pitch, formal) Today, jam sandwiches are mainly consumed by children. Shops do not often sell individual jam sandwiches. == Ingredients and nutrition ==
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 0.0/10; 10.3s, EN.
EN_SKjE_MI17TA_W000005 · in -16.6 dBFS · gain -3.4 dB · emolia-01762
Fatigue Exhaustion ↓  /  Concentrationidentity −0.02 emotion 94 %   proxy_spearman__PXR__T0.50__C0.25__INTERNAL · #12

This chain comes from the proxy rule: the same two-sided test as above, but because Concentration is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Concentration below average — 0.30, lower than 70 % of clips in this corpus — and ends with it at the very top of the corpus at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.63.

At the same time Fatigue Exhaustion goes the other way, from 0.89 (higher than 89 % of clips in this corpus) to 0.33 (lower than 67 % of clips in this corpus), a change of -0.55. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.15, then +0.25, then +0.23 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.63 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.63 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.63, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 31 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.711 before conversion and 0.690 after — it fell by 0.020. Neighbour-to-neighbour the worst pair went 0.644 → 0.608. (The earlier render, with segment 1 left raw, scores 0.564 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.626 in the original and +0.586 after conversion — 94 % of the delta retained, which is essentially all of it. On the other named axis, Fatigue Exhaustion, -0.552 became -0.634.

Quality. Mean predicted overall quality across the segments went 2.59 → 2.95 (+0.37) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.711 → 0.690 -0.020identity cos neighbours 0.644 → 0.608d_b rescored +0.626 → +0.586d_a rescored -0.552 → -0.634d_a mined -0.552d_b mined 0.626min_cos_consec (site) 0.6285min_cos_anchor (site) 0.6285dataset emolialang enspeaker EN_Y9SauDuHviQtotal 30.2schain gain +2.7 dBseam step 2.5 dBcrossfades 150/150/100 ms
Script — 4 chunks, 3 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, fairly smooth, normally alert, fairly steady, moderate pitch range, light breath
(normal-paced, slightly relaxed, some disfluency, casual) Okay, so as far as the Padupi 3, Padupi 3 is one of the, (low mumble) uh, I mean, blue flag beach.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: casual, playful; average recording, some background noise; genuineness 4.1/6; vocal-burst blend 4.0/10; 6.0s, EN.
EN_Y9SauDuHviQ_W000151 · in -16.9 dBFS · gain -3.1 dB · emolia-01222
(normal-paced, slightly relaxed, some disfluency, casual) Before it became a blue flag beach, which I showed in the previous slide,
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; average recording, quiet background; genuineness 2.8/6; vocal-burst blend 2.5/10; 4.1s, EN.
EN_Y9SauDuHviQ_W000152 · in -15.6 dBFS · gain -4.4 dB · emolia-01222
(awe, intoxication altered states of consciousness, interest · normal-paced, neutral tension, frequent disfluency, casual) The total area was like that. I mean, you could see the red (ahem) (ahem) point where the present (ahem) (low mumble) blue flag beach exists around two and a half kilometers and (ahem) the white sandy beaches are no longer there.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, very dark, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as awe, intoxication altered states of consciousness, interest; style: casual, monologue; below-average recording, quiet background; genuineness 5.1/6; vocal-burst blend 9.9/10; 12.9s, EN.
EN_Y9SauDuHviQ_W000153 · in -17.2 dBFS · gain -2.8 dB · emolia-01222
(concentration, anger · brisk, slightly relaxed, some disfluency, monologue) Now, grabbing of coastal wetlands by the state, this is another important (ahem) factor, undermining, I mean, which is response, which is
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as concentration, anger; style: monologue, formal; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 3.3/10; 7.8s, EN.
EN_Y9SauDuHviQ_W000154 · in -18.9 dBFS · gain -1.1 dB · emolia-01222
Sourness ↓  /  Confusionidentity −0.02 emotion 36 %   proxy_spearman__PXR__T0.50__C0.25__INTERNAL · #13

This chain comes from the proxy rule: the same two-sided test as above, but because Confusion is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Confusion below average — 0.34, lower than 66 % of clips in this corpus — and ends with it at the very top of the corpus at 0.91, higher than 91 % of clips in this corpus. That is a total rise of 0.57.

At the same time Sourness goes the other way, from 0.75 (higher than 75 % of clips in this corpus) to 0.24 (lower than 76 % of clips in this corpus), a change of -0.50. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.24, then +0.08, then +0.24 — an uneven climb, but always in the same direction.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.94 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.93 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.94), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 35 s · ko · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.923 before conversion and 0.901 after — it fell by 0.022. Neighbour-to-neighbour the worst pair went 0.923 → 0.901. (The earlier render, with segment 1 left raw, scores 0.833 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Confusion moved +0.567 in the original and +0.202 after conversion — 36 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Sourness, -0.502 became -0.220.

Quality. Mean predicted overall quality across the segments went 3.15 → 3.23 (+0.08) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.923 → 0.901 -0.022identity cos neighbours 0.923 → 0.901d_b rescored +0.567 → +0.202d_a rescored -0.502 → -0.220d_a mined -0.502d_b mined 0.566min_cos_consec (site) 0.9313min_cos_anchor (site) 0.9378dataset emolialang kospeaker KO_qkoobl4yzBUtotal 33.7schain gain +0.1 dBseam step 0.3 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed
(measured, monologue, didactic) 희토류 관련 주 강세에 해인과 수산중공업도 희토류 관련으로 찌라시 올라왔고, 해인은 고가 8%, 수산중공업은 별 반응 없었네요.
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, didactic; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 1.4/10; 10.0s, KO.
KO_qkoobl4yzBU_W000046 · in -17.2 dBFS · gain -2.8 dB · emolia-03061
(measured, monologue, didactic) 태경, BK 역시 히토류 광산을 보유하고 있다는 점이 부각되면서 지랄이 올라왔고, 고가는 5%입니다.
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, didactic; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 3.4/10; 6.9s, KO.
KO_qkoobl4yzBU_W000047 · in -17.7 dBFS · gain -2.3 dB · emolia-03061
(distress · normal-paced, monologue, narration) DIC도 마찬가지로 히토류 관련으로 찌라시 올라왔고, 히토류 제조 방법 특허를 받은 바가 부각되었습니다. 오늘 고가는 5%입니다.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as distress; style: monologue, narration; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 3.1/10; 9.6s, KO.
KO_qkoobl4yzBU_W000048 · in -18.6 dBFS · gain -1.4 dB · emolia-03061
(confusion · measured, monologue, narration) 중국의 히토류 자석 기술 글로벌 점유율은 80에서 90%의 달에 국내 산업계가 긴장하는 가운데
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as confusion; style: monologue, narration; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 2.7/10; 7.8s, KO.
KO_qkoobl4yzBU_W000049 · in -17.3 dBFS · gain -2.7 dB · emolia-03061
Confusion ↓  /  Contemptidentity +0.18 emotion 58 %   proxy_spearman__PXR__T0.50__C0.25__INTERNAL · #14

This chain comes from the proxy rule: the same two-sided test as above, but because Contempt is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Contempt around average — 0.47, lower than 53 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.52.

At the same time Confusion goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.49 (lower than 51 % of clips in this corpus), a change of -0.51. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.23, then +0.03, then +0.20, then +0.06 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.52 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.64 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.52, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 48 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.479 before conversion and 0.661 after — it rose by 0.182. Neighbour-to-neighbour the worst pair went 0.639 → 0.735. (The earlier render, with segment 1 left raw, scores 0.447 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contempt moved +0.522 in the original and +0.304 after conversion — 58 % of the delta retained. On the other named axis, Confusion, -0.512 became -0.660.

Quality. Mean predicted overall quality across the segments went 2.90 → 3.08 (+0.18) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.479 → 0.661 +0.182identity cos neighbours 0.639 → 0.735d_b rescored +0.522 → +0.304d_a rescored -0.512 → -0.660d_a mined -0.512d_b mined 0.522min_cos_consec (site) 0.6411min_cos_anchor (site) 0.5152dataset emolialang enspeaker EN_B00007_S01629total 47.0schain gain +1.9 dBseam step 1.9 dBcrossfades 100/150/150/150 ms
Script — 5 chunks, 4 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · average recording, moderately variable
(confusion, embarrassment, intoxication altered states of consciousness · normal-paced, normally alert, neutral tension, casual) (exhausted groan) (ahem) (ahem) And he's in Hebrew, he's asking me, can I get a roll of film? And I don't turn around, like I don't hear, I don't know what he's saying. (ahem) (low mumble) Say it in (low mumble) Hebrew, say it in Hebrew.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, neutral openness; reads as confusion, embarrassment, intoxication altered states of consciousness; style: casual, conversational; average recording, some background noise; mildly explicit content; genuineness 6.0/6; vocal-burst blend 7.9/10; 11.0s, EN.
EN_B00007_S01629_W000006 · in -18.0 dBFS · gain -2.0 dB · emolia-00407
(confusion, intoxication altered states of consciousness, sexual lust · measured, normally alert, neutral tension, casual) And I don't, I don't turn around and he says again, (ahem) (low mumble) and I turn around and I go, are you talking to me? And he goes, yes. And I said, uh, I (low mumble) don't understand what you're saying. He goes,
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, neutral tension, moderately variable; timbre is neutral-toned, dark, slightly rough, slightly thin; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is mildly negative, slightly dominant, neutral openness; reads as confusion, intoxication altered states of consciousness, sexual lust; style: casual, conversational; average recording, quiet background; genuineness 5.1/6; vocal-burst blend 7.3/10; 13.5s, EN.
EN_B00007_S01629_W000007 · in -19.7 dBFS · gain -0.3 dB · emolia-00407
(amusement, intoxication altered states of consciousness, astonishment surprise · normal-paced, energised, neutral tension, storytelling) And then finally, so he believes me, and he starts with his English, and his jaw is shaking, he's, (ahem) uh, I need, (low mumble) uh, 200, uh, (ahem) speed film.
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as amusement, intoxication altered states of consciousness, astonishment surprise; style: storytelling, playful; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 2.1/10; 8.3s, EN.
EN_B00007_S01629_W000008 · in -17.1 dBFS · gain -2.9 dB · emolia-00407
(impatience and irritability, jealousy and envy, bitterness · brisk, energised, neutral tension, casual) Listen, I worked in that store for two years. I met people from my homeroom class in high school. They didn't ask for a fucking discount.
full caption & clip details
An adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, very rough, thin; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, fairly guarded; reads as impatience and irritability, jealousy and envy, bitterness; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 4.8/6; vocal-burst blend 4.7/10; 7.9s, EN.
EN_B00007_S01629_W000009 · in -18.1 dBFS · gain -1.9 dB · emolia-00407
(contempt, bitterness, impatience and irritability · brisk, energised, tense, casual) They didn't ask for shit. (chuckle) They paid, they said it was great to see you, take care. That's hilarious. It's like just cause you're from Israel and I'm from Israel.
full caption & clip details
An adult masculine voice; delivery is energised, brisk, tense, moderately variable; timbre is slightly cool, slightly bright, slightly rough, thin; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, guarded; reads as contempt, bitterness, impatience and irritability; style: casual, storytelling; average recording, some background noise; mildly explicit content; genuineness 3.3/6; vocal-burst blend 2.3/10; 7.0s, EN.
EN_B00007_S01629_W000010 · in -17.9 dBFS · gain -2.1 dB · emolia-00407
Fatigue Exhaustion ↓  /  Disgustidentity −0.00 emotion 92 %   proxy_spearman__PXR__T0.50__C0.25__INTERNAL · #15

This chain comes from the proxy rule: the same two-sided test as above, but because Disgust is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Disgust below average — 0.33, lower than 67 % of clips in this corpus — and ends with it strongly present at 0.88, higher than 88 % of clips in this corpus. That is a total rise of 0.55.

At the same time Fatigue Exhaustion goes the other way, from 0.85 (higher than 85 % of clips in this corpus) to 0.24 (lower than 76 % of clips in this corpus), a change of -0.61. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.00, then +0.20, then +0.18, then +0.17 — a plateau around step 1, where it barely moves.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.88 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.83 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.88. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 44 s · ko · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.861 before conversion and 0.857 after — it fell by 0.003. Neighbour-to-neighbour the worst pair went 0.816 → 0.838. (The earlier render, with segment 1 left raw, scores 0.760 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disgust moved +0.553 in the original and +0.507 after conversion — 92 % of the delta retained, which is essentially all of it. On the other named axis, Fatigue Exhaustion, -0.609 became -0.670.

Quality. Mean predicted overall quality across the segments went 2.96 → 3.16 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.861 → 0.857 -0.003identity cos neighbours 0.816 → 0.838d_b rescored +0.553 → +0.507d_a rescored -0.609 → -0.670d_a mined -0.609d_b mined 0.553min_cos_consec (site) 0.8308min_cos_anchor (site) 0.8777dataset emolialang kospeaker KO_YSNFaeTzKfAtotal 43.1schain gain -1.4 dBseam step 1.8 dBcrossfades 100/150/150/150 ms
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, average recording, measured, normally alert, slightly relaxed, fairly steady
(somewhat unclear, monologue, conversational) (low mumble) 커피 몇 그릴 가지고 실험한 게 아니고, 현장 포장해서 (low mumble) 구획을 나누어서 실험을 하게 되었습니다.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, conversational; average recording, quiet background; genuineness 2.4/6; vocal-burst blend 4.1/10; 6.0s, KO.
KO_YSNFaeTzKfA_W000033 · in -18.3 dBFS · gain -1.7 dB · emolia-03037
(somewhat unclear, monologue, didactic) 한 구액당 한 250g의 커피나무가 있었구요. 그 다음에 A, B, C는 농도를 달리해서 자닮식 방법을 쓰는 것이고, D는 비격으로 해서 비교를 하게 됩니다.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, didactic; average recording, no background noise; genuineness 1.7/6; vocal-burst blend 2.4/10; 10.6s, KO.
KO_YSNFaeTzKfA_W000034 · in -18.2 dBFS · gain -1.8 dB · emolia-03037
(longing · somewhat unclear, monologue, authoritative) 하와이는 은행이 없었기 때문에 은행을 뺐구요. 약책을 빼고, 자닮오일, 자닮요황,
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as longing; style: monologue, authoritative; average recording, no background noise; genuineness 2.4/6; vocal-burst blend 3.5/10; 5.6s, KO.
KO_YSNFaeTzKfA_W000035 · in -18.3 dBFS · gain -1.7 dB · emolia-03037
(pride · average clarity, monologue, conversational) (ahem) 그 다음에, 가성소다, 어, (wistful sigh) 황토 분말. 이 네 가지의 비율을 최등을 둬가지고, 뭐, ABC 구한을 14일에 한 번 정도씩 방제를 했습니다.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pride; style: monologue, conversational; average recording, quiet background; genuineness 1.4/6; vocal-burst blend 4.8/10; 8.5s, KO.
KO_YSNFaeTzKfA_W000036 · in -16.9 dBFS · gain -3.1 dB · emolia-03037
(somewhat unclear, monologue, didactic) 제가 하와이에서 1년 동안 상주할 수 없었기 때문에 연구보조로 저희 아내와 (low mumble) 딸이 함께 했고요. 어, (low mumble) 그리고 질문 잡고 하셨는데요. 하와이 CGNF 회장인 김장께서 같이 참여하셨습니다.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, didactic; average recording, quiet background; genuineness 2.0/6; vocal-burst blend 5.8/10; 13.2s, KO.
KO_YSNFaeTzKfA_W000037 · in -17.8 dBFS · gain -2.2 dB · emolia-03037
Infatuation ↓  /  Longingidentity +0.08 emotion 7 %   proxy_spearman__PXR__T0.50__C0.25__INTERNAL · #16

This chain comes from the proxy rule: the same two-sided test as above, but because Longing is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Longing below average — 0.34, lower than 66 % of clips in this corpus — and ends with it strongly present at 0.90, higher than 90 % of clips in this corpus. That is a total rise of 0.56.

At the same time Infatuation goes the other way, from 0.77 (higher than 77 % of clips in this corpus) to 0.14 (lower than 86 % of clips in this corpus), a change of -0.63. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.12, then +0.20, then +0.16, then +0.08 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.76 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.76 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.76, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 36 s · snippets

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.676 before conversion and 0.755 after — it rose by 0.079. Neighbour-to-neighbour the worst pair went 0.724 → 0.755. (The earlier render, with segment 1 left raw, scores 0.577 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Longing moved +0.558 in the original and +0.041 after conversion — 7 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Infatuation, -0.615 became -0.352.

Quality. Mean predicted overall quality across the segments went 2.73 → 2.96 (+0.23) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.676 → 0.755 +0.079identity cos neighbours 0.724 → 0.755d_b rescored +0.558 → +0.041d_a rescored -0.615 → -0.352d_a mined -0.631d_b mined 0.557min_cos_consec (site) 0.7573min_cos_anchor (site) 0.7573dataset snippetslang undspeaker batch7_part1_batch7_part1_total 34.8schain gain +3.7 dBseam step 3.0 dBcrossfades 100/150/150/150 ms
Script — 5 chunks, 4 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, normally alert, slightly relaxed, fairly steady
(normal-paced, some disfluency, average clarity, casual) (ahem) Um, and it it's a lot of (low mumble) uh mainstream writers and creators.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: casual, monologue; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 1.5/10; 4.3s.
batch7_part1_batch7_part1_chunk_1058_1_1025065 · in -15.3 dBFS · gain -4.7 dB · snippets-01299
(brisk, little disfluency, clear, formal) (ahem) Uh, the recent news about Eric July and the success of his Rippaverse has got me thinking about how the underground comic scene has changed over the years.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 1.1/10; 7.5s.
batch7_part1_batch7_part1_chunk_1058_1_1025142 · in -15.6 dBFS · gain -4.4 dB · snippets-01299
(normal-paced, some disfluency, average clarity, conversational) (low mumble) Um, well let's get the easy part out of the way first. I I I this has nothing to do with uh a (ahem) CG or any of that kind of stuff and and (ahem) um
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: conversational, casual; good recording, quiet background; genuineness 3.4/6; vocal-burst blend 3.2/10; 8.0s.
batch7_part1_batch7_part1_chunk_1058_1_1025179 · in -16.4 dBFS · gain -3.6 dB · snippets-01299
(contentment, contemplation, elation · normal-paced, some disfluency, average clarity, conversational) these days. That kind of, that kind of rawness to to the work. It's one of the things that I like about Larry King's work and I did the book with him. (ahem) Uh, Jim and I jumped here what a year ago.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as contentment, contemplation, elation; style: conversational, casual; good recording, quiet background; genuineness 4.3/6; vocal-burst blend 4.1/10; 10.1s.
batch7_part1_batch7_part1_chunk_1058_1_1025238 · in -16.8 dBFS · gain -3.2 dB · snippets-01299
(normal-paced, some disfluency, average clarity, casual) you know, a lot of that money went to DC or went to others. Didn't actually go to Alan Moore.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, conversational; good recording, quiet background; genuineness 2.7/6; vocal-burst blend 3.3/10; 5.6s.
batch7_part1_batch7_part1_chunk_1058_1_1025259 · in -17.2 dBFS · gain -2.8 dB · snippets-01299
Infatuation ↓  /  Triumphidentity −0.08 emotion 90 %   proxy_spearman__PXR__T0.50__C0.25__INTERNAL · #17

This chain comes from the proxy rule: the same two-sided test as above, but because Triumph is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Triumph around average — 0.45, lower than 55 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.52.

At the same time Infatuation goes the other way, from 0.83 (higher than 83 % of clips in this corpus) to 0.33 (lower than 67 % of clips in this corpus), a change of -0.50. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.16, then +0.13, then +0.23 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.88 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.92 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.88. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 33 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.871 before conversion and 0.792 after — it fell by 0.079. Neighbour-to-neighbour the worst pair went 0.854 → 0.748. (The earlier render, with segment 1 left raw, scores 0.630 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Triumph moved +0.519 in the original and +0.468 after conversion — 90 % of the delta retained, which is essentially all of it. On the other named axis, Infatuation, -0.504 became -0.407.

Quality. Mean predicted overall quality across the segments went 2.99 → 3.15 (+0.17) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.871 → 0.792 -0.079identity cos neighbours 0.854 → 0.748d_b rescored +0.519 → +0.468d_a rescored -0.504 → -0.407d_a mined -0.504d_b mined 0.519min_cos_consec (site) 0.9186min_cos_anchor (site) 0.8826dataset emolialang enspeaker EN_yl8nD01uB-4total 32.4schain gain +1.7 dBseam step 0.9 dBcrossfades 150/100/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(steady, formal, monologue) The names of the last two clans, the Samakas and the Srinjayas, are also mentioned in the Mahabharata and the Puranas
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.4/10; 6.5s, EN.
EN_yl8nD01uB-4_W000016 · in -14.8 dBFS · gain -5.2 dB · emolia-00810
(steady, formal, authoritative) King Drupada, whose daughter Draupadi was married into the Pandavas, belonged to the Samaka clan
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.2/10; 5.5s, EN.
EN_yl8nD01uB-4_W000017 · in -15.2 dBFS · gain -4.8 dB · emolia-00810
(concentration · fairly steady, newsreading, formal) However, the Mahabharata and the Puranas consider the ruling clan of the northern Panchala as an offshoot of the Bharata clan and Devodassa, Sudas, Srinjaya, Samaka, and Drupada
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: newsreading, formal; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.0/10; 14.0s, EN.
EN_yl8nD01uB-4_W000018 · in -15.1 dBFS · gain -4.9 dB · emolia-00810
(triumph · fairly steady, formal, newsreading) The Panchala Kingdom rose to its highest prominence in the aftermath of the decline and defeat of the Kuru Kingdom by the non-Vedic Salva tribe.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as triumph; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.8/10; 7.0s, EN.
EN_yl8nD01uB-4_W000019 · in -14.6 dBFS · gain -5.4 dB · emolia-00810
Pain ↓  /  Thankfulness Gratitudeidentity −0.13 emotion REVERSED   proxy_spearman__PXR__T0.50__C0.25__INTERNAL · #18

This chain comes from the proxy rule: the same two-sided test as above, but because Thankfulness Gratitude is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Thankfulness Gratitude below average — 0.33, lower than 67 % of clips in this corpus — and ends with it strongly present at 0.84, higher than 84 % of clips in this corpus. That is a total rise of 0.52.

At the same time Pain goes the other way, from 0.86 (higher than 86 % of clips in this corpus) to 0.35 (lower than 65 % of clips in this corpus), a change of -0.51. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.22, then +0.07, then +0.23 — an uneven climb, but always in the same direction.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.76 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.78 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.76, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 32 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.725 before conversion and 0.596 after — it fell by 0.130. Neighbour-to-neighbour the worst pair went 0.695 → 0.562. (The earlier render, with segment 1 left raw, scores 0.540 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

The emotional move did not survive. Re-scored end to end, Thankfulness Gratitude moved +0.518 in the original and -0.249 after conversion — it changed direction. On this chain the corrected audio is not an improvement. On the other named axis, Pain, -0.514 became -0.502.

Quality. Mean predicted overall quality across the segments went 2.29 → 2.92 (+0.63) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.725 → 0.596 -0.130identity cos neighbours 0.695 → 0.562d_b rescored +0.518 → -0.249d_a rescored -0.514 → -0.502d_a mined -0.514d_b mined 0.518min_cos_consec (site) 0.7751min_cos_anchor (site) 0.7576dataset emolialang enspeaker EN_w8NvhHj82gktotal 30.6schain gain +3.2 dBseam step 2.7 dBcrossfades 100/100/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · neutral-toned, slightly bright, fairly smooth, good recording, no background noise
(normal-paced, normally alert, slightly relaxed, casual) What I really care about for this class is pretty much right here. It's the files section.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: casual, dramatic; good recording, no background noise; genuineness 1.6/6; vocal-burst blend 0.0/10; 5.9s, EN.
EN_w8NvhHj82gk_W000023 · in -15.1 dBFS · gain -5.0 dB · emolia-01294
(contentment, pride, contempt · normal-paced, normally alert, slightly relaxed, dramatic) You don't really need to do anything about going into cPanel after the first time. We just need to make sure that you have everything ready so that you can FTP, which stands for File Transfer Protocol.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, slightly bright, fairly smooth, full; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as contentment, pride, contempt; style: dramatic, casual; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.6/10; 11.2s, EN.
EN_w8NvhHj82gk_W000024 · in -13.6 dBFS · gain -6.4 dB · emolia-01294
(malevolence malice, sexual lust, sourness · slow, energised, slightly tense, casual) That will allow you to upload your files, your images, your HTML pages to your web server. But it's a good idea to know how to get in here.
full caption & clip details
A young adult feminine voice; delivery is energised, slow, slightly tense, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, full; clear, some disfluency, wide pitch range, normal breath; affect is mildly positive, slightly dominant, neutral openness; reads as malevolence malice, sexual lust, sourness; style: casual, dramatic; good recording, no background noise; genuineness 1.4/6; vocal-burst blend 1.1/10; 10.5s, EN.
EN_w8NvhHj82gk_W000025 · in -14.2 dBFS · gain -5.8 dB · emolia-01294
(normal-paced, normally alert, slightly relaxed, casual) So we're going to largely work within the FTP server.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: casual, monologue; good recording, no background noise; genuineness 2.0/6; vocal-burst blend 3.2/10; 3.5s, EN.
EN_w8NvhHj82gk_W000026 · in -12.1 dBFS · gain -7.9 dB · emolia-01294
Infatuation ↓  /  Thankfulness Gratitudeidentity −0.06 emotion 125 %   proxy_spearman__PXR__T0.50__C0.25__INTERNAL · #19

This chain comes from the proxy rule: the same two-sided test as above, but because Thankfulness Gratitude is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Thankfulness Gratitude below average — 0.33, lower than 67 % of clips in this corpus — and ends with it at the very top of the corpus at 0.90, higher than 90 % of clips in this corpus. That is a total rise of 0.58.

At the same time Infatuation goes the other way, from 0.91 (higher than 91 % of clips in this corpus) to 0.39 (lower than 61 % of clips in this corpus), a change of -0.52. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.03, then +0.24, then +0.12, then +0.19 — a plateau around step 1, where it barely moves.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.90 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.87 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.90), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 52 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.918 before conversion and 0.857 after — it fell by 0.061. Neighbour-to-neighbour the worst pair went 0.882 → 0.844. (The earlier render, with segment 1 left raw, scores 0.509 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Thankfulness Gratitude moved +0.576 in the original and +0.721 after conversion — 125 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Infatuation, -0.523 became -0.664.

Quality. Mean predicted overall quality across the segments went 3.06 → 3.21 (+0.15) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.918 → 0.857 -0.061identity cos neighbours 0.882 → 0.844d_b rescored +0.576 → +0.721d_a rescored -0.523 → -0.664d_a mined -0.523d_b mined 0.576min_cos_consec (site) 0.8695min_cos_anchor (site) 0.9011dataset emolialang enspeaker EN_sk-csK-eCZ4total 50.9schain gain +1.0 dBseam step 0.8 dBcrossfades 150/150/100/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(infatuation · steady, newsreading, formal) In Chef Robert Carrier's recipe for it, the base is made from yeast pastry rather than often used shortcrust pastry, because the yeast pastry
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as infatuation; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 11.5s, EN.
EN_sk-csK-eCZ4_W000037 · in -16.4 dBFS · gain -3.6 dB · emolia-02611
(disgust · steady, newsreading, formal) In Italy, plum cake is known by the English name, baked in an oven using dried fruit and often yogurt.The Polish version of plum cake, which also uses fresh fruit, is known as Plasek z śloukami
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disgust; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 10.7s, EN.
EN_sk-csK-eCZ4_W000038 · in -14.6 dBFS · gain -5.4 dB · emolia-02611
(fairly steady, formal, newsreading) In India plum cake has been served around the time of the Christmas holiday season, and may have additional ingredients such as rum or brandy added
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, newsreading; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.0/10; 7.4s, EN.
EN_sk-csK-eCZ4_W000040 · in -14.0 dBFS · gain -6.0 dB · emolia-02611
(steady, newsreading, formal) Plum cake in the United States originated with the English settlers and was prepared in the English style in sizes ranging from small, such as for parties in celebration of Twelfth Night and Christmas, to large, such as for weddings
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: newsreading, formal; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.0/10; 12.6s, EN.
EN_sk-csK-eCZ4_W000042 · in -15.0 dBFS · gain -5.0 dB · emolia-02611
(thankfulness gratitude · steady, formal, newsreading) This original fruitcake version of plum cake in the United States has been referred to as a reigning,
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as thankfulness gratitude; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 9.5s, EN.
EN_sk-csK-eCZ4_W000043 · in -14.3 dBFS · gain -5.7 dB · emolia-02611
Fatigue Exhaustion ↓  /  Concentrationidentity −0.04 emotion 137 %   proxy_spearman__PXR__T0.50__C0.25__INTERNAL · #20

This chain comes from the proxy rule: the same two-sided test as above, but because Concentration is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Concentration below average — 0.29, lower than 71 % of clips in this corpus — and ends with it strongly present at 0.80, higher than 80 % of clips in this corpus. That is a total rise of 0.51.

At the same time Fatigue Exhaustion goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.40 (lower than 60 % of clips in this corpus), a change of -0.56. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.17, then +0.17, then +0.17, then +0.00 — a plateau around step 4, where it barely moves.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.84 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.84 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.84. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 37 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.632 before conversion and 0.588 after — it fell by 0.044. Neighbour-to-neighbour the worst pair went 0.769 → 0.766. (The earlier render, with segment 1 left raw, scores 0.493 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.509 in the original and +0.697 after conversion — 137 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Fatigue Exhaustion, -0.558 became -0.282.

Quality. Mean predicted overall quality across the segments went 2.81 → 2.87 (+0.07) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.632 → 0.588 -0.044identity cos neighbours 0.769 → 0.766d_b rescored +0.509 → +0.697d_a rescored -0.558 → -0.282d_a mined -0.558d_b mined 0.508min_cos_consec (site) 0.8429min_cos_anchor (site) 0.8429dataset emolialang enspeaker EN_B00042_S00699total 35.4schain gain +2.1 dBseam step 1.6 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, balanced body, slightly relaxed, light breath
(fatigue exhaustion, doubt · normal-paced, normally alert, fairly steady, casual) Plunging perhaps may be too strong over, but that's been sinking for some month now.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; average clarity, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fatigue exhaustion, doubt; style: casual, monologue; average recording, no background noise; genuineness 2.6/6; vocal-burst blend 0.8/10; 3.7s, EN.
EN_B00042_S00699_W000048 · in -35.0 dBFS · gain +15.0 dB · emolia-01067
(bitterness, thankfulness gratitude, distress · normal-paced, normally alert, fairly steady, formal) Because America has been living beyond her means, borrowing two billion a day from foreign nations.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as bitterness, thankfulness gratitude, distress; style: formal, monologue; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 0.9/10; 4.5s, EN.
EN_B00042_S00699_W000049 · in -29.3 dBFS · gain +9.3 dB · emolia-01067
(measured, normally alert, fairly steady, narration) The dollar has plummeted more in Bush's term than during any comparable period of US history. A sinking dollar means a poorer nation.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, formal; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 0.1/10; 7.9s, EN.
EN_B00042_S00699_W000050 · in -29.9 dBFS · gain +9.8 dB · emolia-01067
(measured, normally alert, steady, monologue) And a superpower with a sinking currency is a contradiction in terms. Close quote. Unscramble that one.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; good recording, no background noise; genuineness 1.3/6; vocal-burst blend 0.5/10; 6.9s, EN.
EN_B00042_S00699_W000051 · in -32.8 dBFS · gain +12.8 dB · emolia-01067
(measured, subdued, fairly steady, whispered) John, let me give you one more, one other statistic. (ahem) Uh, during the last seven years, the Bush presidency, (contented sigh) household debt has almost tripled. You take total household debt and total federal debt and combine them.
full caption & clip details
An elderly masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; slurred, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: whispered, didactic; average recording, quiet background; genuineness 2.0/6; vocal-burst blend 0.5/10; 13.4s, EN.
EN_B00042_S00699_W000052 · in -33.0 dBFS · gain +13.0 dB · emolia-01067