proxy_spearman__PXR__T0.40__C0.25__INTERNAL — voice-corrected

Manifest tier. proxy_spearman, rule PXR, T=0.4, step cap 0.25. Population 12,053 chains (135 h) over 6 corpora. The SHAREABLE variant of this tier (podcast and evasnippets excluded) holds 10,156.

This is the voice-conversion-corrected version of this tier. The original, uncorrected page is still there and unchanged: t_proxy_spearman__PXR__T0.40__C0.25__INTERNAL.html. Every card below carries both renders so you can switch between them without leaving the page. What was done, and what it measured →
20chains converted
64segments re-voiced
0.760 → 0.736median worst-to-anchor identity cosine
79 %median emotion delta retained
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken.

(underlined, plain · delivery, style) — the tag before the words. Emotions first, then how it is delivered. Underlined descriptors are the ones that change across this chain — anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words are a real non-speech sound, printed where it happens.

The full generated caption for any clip is under “full caption & clip details”. Its perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.

The tags and captions describe the ORIGINAL clips — they are the corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the converted audio are the ones printed in each card's own paragraph and chips.
Doubt ↓  /  Infatuationidentity −0.01 emotion 80 %   proxy_spearman__PXR__T0.40__C0.25__INTERNAL · #1

This chain comes from the proxy rule: the same two-sided test as above, but because Infatuation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Infatuation below average — 0.30, lower than 70 % of clips in this corpus — and ends with it strongly present at 0.80, higher than 80 % of clips in this corpus. That is a total rise of 0.50.

At the same time Doubt goes the other way, from 0.90 (higher than 90 % of clips in this corpus) to 0.37 (lower than 63 % of clips in this corpus), a change of -0.54. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.03, then +0.21, then +0.16, then +0.10 — a plateau around step 1, where it barely moves.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.83 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.87 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.83. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 51 s · zh · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.780 before conversion and 0.771 after — it fell by 0.009. Neighbour-to-neighbour the worst pair went 0.707 → 0.750. (The earlier render, with segment 1 left raw, scores 0.755 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Infatuation moved +0.496 in the original and +0.399 after conversion — 80 % of the delta retained, which is most of it. On the other named axis, Doubt, -0.536 became -0.220.

Quality. Mean predicted overall quality across the segments went 3.01 → 3.18 (+0.17) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.780 → 0.771 -0.009identity cos neighbours 0.707 → 0.750d_b rescored +0.496 → +0.399d_a rescored -0.536 → -0.220d_a mined -0.535d_b mined 0.496min_cos_consec (site) 0.8684min_cos_anchor (site) 0.8270dataset emolialang zhspeaker ZH_B00058_S05719total 49.2schain gain +1.5 dBseam step 1.0 dBcrossfades 150/150/150/100 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, no background noise, normally alert, slightly relaxed, fairly steady, no disfluency
(doubt · fast, storytelling, monologue) 未来你想要预制菜赶出我们的生活,这可是实在不现实。第一就是现代人没时间。
full caption & clip details
A young adult feminine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as doubt; style: storytelling, monologue; average recording, no background noise; genuineness 1.1/6; vocal-burst blend 4.3/10; 6.5s, ZH.
ZH_B00058_S05719_W000019 · in -18.3 dBFS · gain -1.7 dB · emolia-03861
(fast, dramatic, storytelling) 再一个就是资本的理论,就是快速,谁能够快速复制,谁就占住了资本市场。
full caption & clip details
A young adult feminine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: dramatic, storytelling; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 3.8/10; 6.2s, ZH.
ZH_B00058_S05719_W000020 · in -18.7 dBFS · gain -1.3 dB · emolia-03861
(thankfulness gratitude · normal-paced, monologue, formal) 而豫制菜恰好站到了这一个风口浪尖上,既有年轻人的消费需求,又有资本的运作。高速发展的一个需求。所以预制菜未来将成为我们生活中不可或缺的一部分。
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as thankfulness gratitude; style: monologue, formal; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 3.2/10; 12.9s, ZH.
ZH_B00058_S05719_W000021 · in -18.6 dBFS · gain -1.4 dB · emolia-03861
(thankfulness gratitude · normal-paced, monologue, formal) 只不过这个春节还是主打一个团圆。既然是团圆,一家人整整齐齐的聚在一起,聊天也好,畅谈也好,更多的不妨围在厨房一起包一顿饺子,一起啊做一做咱们黄陂传统的炸圆子酿藕汤这样的一个黄陂三鲜。
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as thankfulness gratitude; style: monologue, formal; average recording, no background noise; genuineness 0.9/6; vocal-burst blend 3.9/10; 18.5s, ZH.
ZH_B00058_S05719_W000022 · in -17.8 dBFS · gain -2.2 dB · emolia-03861
(measured, narration, storytelling) 玉制菜虽好,但是也不要全篇覆盖该有的,传统餐饮还是要有的。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, storytelling; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 4.0/10; 6.0s, ZH.
ZH_B00058_S05719_W000023 · in -19.9 dBFS · gain -0.1 dB · emolia-03861
Astonishment Surprise ↓  /  Confusionidentity −0.01 emotion 96 %   proxy_spearman__PXR__T0.40__C0.25__INTERNAL · #2

This chain comes from the proxy rule: the same two-sided test as above, but because Confusion is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Confusion around average — 0.49, lower than 51 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.49.

At the same time Astonishment Surprise goes the other way, from 0.72 (higher than 72 % of clips in this corpus) to 0.23 (lower than 78 % of clips in this corpus), a change of -0.50. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.23, then +0.16, then +0.09 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.90 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.90 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.90. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 34 s · zh · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.838 before conversion and 0.833 after — it fell by 0.005. Neighbour-to-neighbour the worst pair went 0.866 → 0.853. (The earlier render, with segment 1 left raw, scores 0.795 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Confusion moved +0.488 in the original and +0.469 after conversion — 96 % of the delta retained, which is essentially all of it. On the other named axis, Astonishment Surprise, -0.498 became -0.165.

Quality. Mean predicted overall quality across the segments went 2.98 → 3.21 (+0.22) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.838 → 0.833 -0.005identity cos neighbours 0.866 → 0.853d_b rescored +0.488 → +0.469d_a rescored -0.498 → -0.165d_a mined -0.498d_b mined 0.488min_cos_consec (site) 0.8965min_cos_anchor (site) 0.8964dataset emolialang zhspeaker ZH_B00017_S02767total 32.5schain gain +0.7 dBseam step 1.0 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · neutral-toned, fairly smooth, average recording, no background noise, measured, normally alert, slightly relaxed
(fairly steady, no disfluency, clear, didactic) 外部环境的突发事件等。因此,企业需要建立一个有效的监控机制。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, formal; average recording, no background noise; genuineness 0.6/6; vocal-burst blend 1.1/10; 6.0s, ZH.
ZH_B00017_S02767_W000058 · in -17.7 dBFS · gain -2.3 dB · emolia-03449
(steady, no disfluency, clear, monologue) 已实时了解战略执行的情况,这个监控机制应该包括。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, didactic; average recording, no background noise; genuineness 0.9/6; vocal-burst blend 0.4/10; 6.3s, ZH.
ZH_B00017_S02767_W000059 · in -17.1 dBFS · gain -2.9 dB · emolia-03449
(thankfulness gratitude · fairly steady, some disfluency, somewhat unclear, didactic) 定期的战略执行情况报告以及实时的数据反馈和分析系统。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, thin; somewhat unclear, some disfluency, narrow pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; reads as thankfulness gratitude; style: didactic, monologue; average recording, no background noise; genuineness 1.3/6; vocal-burst blend 1.2/10; 8.4s, ZH.
ZH_B00017_S02767_W000060 · in -17.0 dBFS · gain -3.0 dB · emolia-03449
(confusion, doubt, fatigue exhaustion · fairly steady, some disfluency, clear, didactic) 从而及时采取措施进行改正或调整。同时,根据实际执行情况,可能需要对战略进行实施的调整。这是因为。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as confusion, doubt, fatigue exhaustion; style: didactic, monologue; average recording, no background noise; genuineness 1.2/6; vocal-burst blend 0.4/10; 12.5s, ZH.
ZH_B00017_S02767_W000061 · in -17.1 dBFS · gain -2.9 dB · emolia-03449
Emotional Numbness ↓  /  Malevolence Maliceidentity −0.06 emotion 75 %   proxy_spearman__PXR__T0.40__C0.25__INTERNAL · #3

This chain comes from the proxy rule: the same two-sided test as above, but because Malevolence Malice is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Malevolence Malice around average — 0.47, lower than 53 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.52.

At the same time Emotional Numbness goes the other way, from 0.87 (higher than 87 % of clips in this corpus) to 0.23 (lower than 77 % of clips in this corpus), a change of -0.65. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.05, then +0.23, then +0.24 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.81 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.81 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.81. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 45 s · zh · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.741 before conversion and 0.678 after — it fell by 0.062. Neighbour-to-neighbour the worst pair went 0.689 → 0.764. (The earlier render, with segment 1 left raw, scores 0.666 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Malevolence Malice moved +0.521 in the original and +0.393 after conversion — 75 % of the delta retained, which is most of it. On the other named axis, Emotional Numbness, -0.646 became -0.262.

Quality. Mean predicted overall quality across the segments went 2.85 → 3.11 (+0.26) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.741 → 0.678 -0.062identity cos neighbours 0.689 → 0.764d_b rescored +0.521 → +0.393d_a rescored -0.646 → -0.262d_a mined -0.646d_b mined 0.521min_cos_consec (site) 0.8091min_cos_anchor (site) 0.8091dataset emolialang zhspeaker ZH_B00016_S01299total 44.1schain gain +2.2 dBseam step 0.7 dBcrossfades 150/100/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normally alert, slightly relaxed, fairly steady, some disfluency
(normal-paced, average clarity, formal, conversational) 我们领导人提出正能量这三个字真的非常非常有厉害。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, conversational; good recording, no background noise; genuineness 2.6/6; vocal-burst blend 3.6/10; 3.8s, ZH.
ZH_B00016_S01299_W000030 · in -21.4 dBFS · gain +1.4 dB · emolia-03441
(sexual lust · measured, clear, monologue, authoritative) 你永远要跟正能量的为为伍,你跟正能量人在一起,你会更加正能量。你跟负能量在一起,你会更加负能量。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as sexual lust; style: monologue, authoritative; average recording, no background noise; genuineness 2.3/6; vocal-burst blend 4.9/10; 5.7s, ZH.
ZH_B00016_S01299_W000031 · in -18.9 dBFS · gain -1.1 dB · emolia-03441
(triumph, jealousy and envy, pride · measured, somewhat unclear, monologue, authoritative) 所以我真的觉得很奇怪啊,做好事没人议论一件坏事,所有人全部吸过去。哇,而且各自从兜里面掏出自己的粪,在那里闻你闻闻我的我闻闻你的多恶心啊。虽然这个案例很恶心啊,但是他真的能够表达很能表达我们的观点。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as triumph, jealousy and envy, pride; style: monologue, authoritative; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 8.1/10; 17.0s, ZH.
ZH_B00016_S01299_W000032 · in -18.4 dBFS · gain -1.6 dB · emolia-03441
(malevolence malice, bitterness, triumph · normal-paced, average clarity, storytelling, monologue) 真的,你拉完屎以后已经臭了,你一下了,你从洗手间出来,把那个屎冲走就不臭了。结束了非把屎揣着揣到兜里面,天天见人就给别人闻,然后别人也闻上瘾了,喜欢上这个味儿了,然后自己拉了也开始装到口袋,每一个人这个口袋里装的全是屎。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as malevolence malice, bitterness, triumph; style: storytelling, monologue; average recording, quiet background; genuineness 2.4/6; vocal-burst blend 10.0/10; 18.2s, ZH.
ZH_B00016_S01299_W000033 · in -16.8 dBFS · gain -3.2 dB · emolia-03441
Contemplation ↓  /  Triumphidentity −0.05 emotion 85 %   proxy_spearman__PXR__T0.40__C0.25__INTERNAL · #4

This chain comes from the proxy rule: the same two-sided test as above, but because Triumph is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Triumph below average — 0.38, lower than 62 % of clips in this corpus — and ends with it strongly present at 0.84, higher than 84 % of clips in this corpus. That is a total rise of 0.46.

At the same time Contemplation goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.43 (lower than 57 % of clips in this corpus), a change of -0.56. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.13, then +0.19, then +0.22, then -0.09 — not a clean run: step 4 moves back the other way by 0.09 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.91 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.92 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.91), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 51 s · zh · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.852 before conversion and 0.807 after — it fell by 0.046. Neighbour-to-neighbour the worst pair went 0.881 → 0.806. (The earlier render, with segment 1 left raw, scores 0.686 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Triumph moved +0.457 in the original and +0.387 after conversion — 85 % of the delta retained, which is most of it. On the other named axis, Contemplation, -0.555 became -0.605.

Quality. Mean predicted overall quality across the segments went 2.98 → 3.15 (+0.17) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.852 → 0.807 -0.046identity cos neighbours 0.881 → 0.806d_b rescored +0.457 → +0.387d_a rescored -0.555 → -0.605d_a mined -0.556d_b mined 0.457min_cos_consec (site) 0.9235min_cos_anchor (site) 0.9095dataset emolialang zhspeaker ZH_B00061_S08049total 49.9schain gain +0.2 dBseam step 1.1 dBcrossfades 150/150/100/150 ms
Script — 5 chunks, 5 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, fairly smooth, average recording, quiet background
(contemplation, longing, thankfulness gratitude · measured, subdued, relaxed, monologue) (wistful sigh) (low mumble) (ahem) (ahem) (ahem) 这样的话我们就得到了一个干净的东西啊,还有一些我看看啊,还有这个面组,还有这个组挺烦人的。我们连组也不要啊,加一个组组删除删除组,把所有的主都删除掉啊。
full caption & clip details
An adult masculine voice; delivery is subdued, measured, relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as contemplation, longing, thankfulness gratitude; style: monologue, casual; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 5.8/10; 15.3s, ZH.
ZH_B00061_S08049_W000044 · in -18.4 dBFS · gain -1.6 dB · emolia-03891
(doubt, confusion, contemplation · measured, very low-energy, relaxed, monologue) (ahem) 这样的话我们就会得到一个非常干净的水水点啊。行,有了这一步以后呢,我们继续来做啊,成,那么现在这个就干净了啊,然后接下来VDB. (surprised gasp) (resonant hum) (ahem)
full caption & clip details
An elderly masculine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is neutral-toned, dark, fairly smooth, slightly thin; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly negative, submissive, neutral openness; reads as doubt, confusion, contemplation; style: monologue, whispered; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 5.7/10; 15.2s, ZH.
ZH_B00061_S08049_W000045 · in -20.2 dBFS · gain +0.2 dB · emolia-03891
(confusion · fast, normally alert, slightly relaxed, conversational) (ahem) 这种这个SDF的啊,我们要这个SDF的表面呃,然后呢。 (ahem)
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; slurred, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as confusion; style: conversational, casual; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 3.1/10; 5.3s, ZH.
ZH_B00061_S08049_W000046 · in -18.9 dBFS · gain -1.1 dB · emolia-03891
(triumph · measured, normally alert, slightly relaxed, monologue) 他这边呢会按照这边的这个peace (ahem) scale属性,也就是这边的啊这个零点一二这个大小啊进行一个这个。 (surprised gasp)
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as triumph; style: monologue, didactic; average recording, quiet background; genuineness 3.3/6; vocal-burst blend 5.7/10; 9.0s, ZH.
ZH_B00061_S08049_W000047 · in -17.2 dBFS · gain -2.8 dB · emolia-03891
(fast, normally alert, slightly relaxed, monologue) (low mumble) 然后呢,这边的这个点的这个精度在这里去控制啊,一开始不用太高啊。
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; slurred, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, didactic; average recording, quiet background; genuineness 3.3/6; vocal-burst blend 4.3/10; 5.8s, ZH.
ZH_B00061_S08049_W000048 · in -16.9 dBFS · gain -3.1 dB · emolia-03891
Confusion ↓  /  Fatigue Exhaustionidentity +0.05 emotion REVERSED   proxy_spearman__PXR__T0.40__C0.25__INTERNAL · #5

This chain comes from the proxy rule: the same two-sided test as above, but because Fatigue Exhaustion is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Fatigue Exhaustion around average — 0.45, lower than 55 % of clips in this corpus — and ends with it strongly present at 0.90, higher than 90 % of clips in this corpus. That is a total rise of 0.45.

At the same time Confusion goes the other way, from 0.94 (higher than 94 % of clips in this corpus) to 0.49 (lower than 51 % of clips in this corpus), a change of -0.45. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.20, then +0.19, then -0.13, then +0.19 — not a clean run: step 3 moves back the other way by 0.13 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.72 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.70 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.72, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 33 s · zh · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.602 before conversion and 0.649 after — it rose by 0.046. Neighbour-to-neighbour the worst pair went 0.665 → 0.747. (The earlier render, with segment 1 left raw, scores 0.620 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

The emotional move did not survive. Re-scored end to end, Fatigue Exhaustion moved +0.452 in the original and -0.120 after conversion — it changed direction. On this chain the corrected audio is not an improvement. On the other named axis, Confusion, -0.451 became -0.632.

Quality. Mean predicted overall quality across the segments went 3.01 → 3.22 (+0.21) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.602 → 0.649 +0.046identity cos neighbours 0.665 → 0.747d_b rescored +0.452 → -0.120d_a rescored -0.451 → -0.632d_a mined -0.451d_b mined 0.452min_cos_consec (site) 0.7031min_cos_anchor (site) 0.7173dataset emolialang zhspeaker ZH_B00053_S04172total 32.1schain gain +0.5 dBseam step 2.2 dBcrossfades 150/100/100/150 ms
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: an elderly masculine voice · balanced body, measured, fairly steady
(confusion, doubt · very low-energy, relaxed, frequent disfluency, monologue) (ahem) 明白了,齐王是因国力匮乏,故不得不仰仗和容忍这种穷凶极恶之徒啊,唉,好为他办事啊,压缝到仰仗他们的人是田丹。
full caption & clip details
An elderly masculine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is slightly warm, slightly dark, slightly rough, balanced body; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly negative, neutral stance, neutral openness; reads as confusion, doubt; style: monologue, narration; average recording, quiet background; genuineness 2.6/6; vocal-burst blend 5.5/10; 15.1s, ZH.
ZH_B00053_S04172_W000024 · in -13.6 dBFS · gain -6.4 dB · emolia-03804
(confusion · normally alert, slightly relaxed, no disfluency, formal) 我们一直怀疑,田丹和敖卫谋是同族的异姓兄弟。
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as confusion; style: formal, didactic; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 2.1/10; 4.4s, ZH.
ZH_B00053_S04172_W000025 · in -13.8 dBFS · gain -6.2 dB · emolia-03804
(very low-energy, slightly relaxed, frequent disfluency, ASMR) 若是此人亲来,我们将非常危险。
full caption & clip details
A child feminine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is neutral-toned, dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; style: ASMR, monologue; average recording, no background noise; no dominant emotion; genuineness 3.1/6; vocal-burst blend 5.1/10; 3.4s, ZH.
ZH_B00053_S04172_W000026 · in -14.1 dBFS · gain -5.9 dB · emolia-03804
(normally alert, slightly relaxed, some disfluency, conversational) 雅儿情愿自尽,也不肯落入他的手里。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: conversational, storytelling; average recording, no background noise; genuineness 2.4/6; vocal-burst blend 3.3/10; 3.7s, ZH.
ZH_B00053_S04172_W000027 · in -15.0 dBFS · gain -5.0 dB · emolia-03804
(normally alert, slightly relaxed, no disfluency, monologue) 向少龙听的是,肉跳心惊安慰他一番后,倪夫人忽然来访。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; slurred, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, narration; average recording, no background noise; genuineness 1.1/6; vocal-burst blend 2.5/10; 6.1s, ZH.
ZH_B00053_S04172_W000028 · in -14.4 dBFS · gain -5.6 dB · emolia-03804
Fatigue Exhaustion ↓  /  Hope Enthusiasm Optimismidentity +0.00 emotion 82 %   proxy_spearman__PXR__T0.40__C0.25__INTERNAL · #6

This chain comes from the proxy rule: the same two-sided test as above, but because Hope Enthusiasm Optimism is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Hope Enthusiasm Optimism around average — 0.51, right about the corpus median — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.43.

At the same time Fatigue Exhaustion goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.52 (higher than 52 % of clips in this corpus), a change of -0.47. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.20, then +0.23 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.81 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.86 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.81. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 31 s · en · podcast

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.787 before conversion and 0.788 after — it rose by 0.001. Neighbour-to-neighbour the worst pair went 0.842 → 0.779. (The earlier render, with segment 1 left raw, scores 0.689 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Hope Enthusiasm Optimism moved +0.428 in the original and +0.351 after conversion — 82 % of the delta retained, which is most of it. On the other named axis, Fatigue Exhaustion, -0.468 became -0.241.

Quality. Mean predicted overall quality across the segments went 2.79 → 3.12 (+0.33) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.787 → 0.788 +0.001identity cos neighbours 0.842 → 0.779d_b rescored +0.428 → +0.351d_a rescored -0.468 → -0.241d_a mined -0.468d_b mined 0.433min_cos_consec (site) 0.8585min_cos_anchor (site) 0.8076dataset podcastlang enspeaker 779920total 30.0schain gain +3.7 dBseam step 2.4 dBcrossfades 150/150 ms
Script — 3 chunks, 3 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, slightly rough
(fatigue exhaustion, disappointment, helplessness · measured, very low-energy, relaxed, casual) what you pay sixty Australian dollars for that speed (exhausted groan) uh yeah unfortunately we have to s and you know uh (contented sigh) not sure what to say but yep. If you
full caption & clip details
An adult masculine voice; delivery is very low-energy, measured, relaxed, moderately variable; timbre is neutral-toned, dark, slightly rough, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly negative, slightly submissive, neutral openness; reads as fatigue exhaustion, disappointment, helplessness; style: casual, ASMR; below-average recording, quiet background; genuineness 4.5/6; vocal-burst blend 3.3/10; 13.3s, EN.
779920_00138928 · in -29.2 dBFS · gain +9.2 dB · podcast-00834
(intoxication altered states of consciousness, embarrassment, sexual lust · normal-paced, normally alert, neutral tension, casual) come here for the internet. Just come here for the sun the surf the beach and the chicks but if you want to do any browsing on internet, (low mumble) uh don't bother. Stay in America or stay in Canada.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as intoxication altered states of consciousness, embarrassment, sexual lust; style: casual, monologue; average recording, some background noise; genuineness 4.6/6; vocal-burst blend 6.8/10; 10.2s, EN.
779920_00140424 · in -25.4 dBFS · gain +5.4 dB · podcast-03500
(hope enthusiasm optimism, jealousy and envy · normal-paced, normally alert, neutral tension, casual) Come (ahem) visit me when I'm living there. Yeah. When Lee goes to Canada, everyone listening is all invited to Lee's place to enjoy fast speed internet.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as hope enthusiasm optimism, jealousy and envy; style: casual, conversational; below-average recording, quiet background; genuineness 4.5/6; vocal-burst blend 5.5/10; 6.9s, EN.
779920_00141536 · in -23.1 dBFS · gain +3.1 dB · podcast-03342
Pride ↓  /  Distressidentity −0.01 emotion 91 %   proxy_spearman__PXR__T0.40__C0.25__INTERNAL · #7

This chain comes from the proxy rule: the same two-sided test as above, but because Distress is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Distress around average — 0.44, lower than 56 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.55.

At the same time Pride goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.53 (higher than 53 % of clips in this corpus), a change of -0.46. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.00, then +0.55 — a slow start, with most of the change arriving in the final step.

The largest step is 0.55, which is above the 0.25 cap the strict rule would impose — worth knowing when judging how gradual it sounds.

Same speaker? No similarity score is available here — the eurospeech clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 41 s · it · eurospeech

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.908 before conversion and 0.896 after — it fell by 0.013. Neighbour-to-neighbour the worst pair went 0.908 → 0.881. (The earlier render, with segment 1 left raw, scores 0.766 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Distress moved +0.548 in the original and +0.500 after conversion — 91 % of the delta retained, which is essentially all of it. On the other named axis, Pride, -0.744 became -0.852.

Quality. Mean predicted overall quality across the segments went 2.93 → 3.16 (+0.23) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.908 → 0.896 -0.013identity cos neighbours 0.908 → 0.881d_b rescored +0.548 → +0.500d_a rescored -0.744 → -0.852d_a mined -0.458d_b mined 0.548min_cos_consec (site) —min_cos_anchor (site) —dataset eurospeechlang itspeaker italy_18_365total 40.0schain gain +3.8 dBseam step 0.4 dBcrossfades 150/150 ms
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: a child feminine voice · slightly bright, fairly smooth, average recording, quiet background, fast, normally alert, some disfluency, wide pitch range
(pride · slightly relaxed, moderately variable, somewhat unclear, monologue) (ahem) Certo, ci sono altri fattori di rischio; uno fra tutti, (ahem) vale la pena sottolinearlo, è stato il tema del rincaro delle materie prime, dei semilavorati
full caption & clip details
A child feminine voice; delivery is normally alert, fast, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; somewhat unclear, some disfluency, wide pitch range, light breath; affect is positive, neutral stance, slightly guarded; reads as pride; style: monologue, casual; average recording, quiet background; genuineness 2.3/6; vocal-burst blend 3.8/10; 11.8s, IT.
italy_18_365_13210320_13222128 · in -37.3 dBFS · gain +17.3 dB · eurospeech-01774
(relief, concentration · neutral tension, moderately variable, somewhat unclear, casual) e dei prodotti, ma anche e soprattutto del settore energetico, con particolare riferimento al gas. Nelle previsioni della NADEF l'inflazione ha un rialzo nei primi tre mesi all'1,2 e, su base annua, all'1,5
full caption & clip details
A child feminine voice; delivery is normally alert, fast, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; somewhat unclear, some disfluency, wide pitch range, heavy breath; affect is positive, slightly dominant, slightly guarded; reads as relief, concentration; style: casual, dramatic; average recording, quiet background; genuineness 3.3/6; vocal-burst blend 7.5/10; 18.0s, IT.
italy_18_365_13222128_13240128 · in -38.1 dBFS · gain +18.1 dB · eurospeech-01774
(distress, disappointment, thankfulness gratitude · slightly relaxed, fairly steady, average clarity, dramatic) fa la dichiarazione del presidente Draghi rispetto a un centro di stoccaggio europeo del gas; mi sembra che tali soluzioni vadano nella direzione giusta per (ahem) tutelare
full caption & clip details
A young adult feminine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is slightly cool, slightly bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as distress, disappointment, thankfulness gratitude; style: dramatic, monologue; average recording, quiet background; genuineness 1.7/6; vocal-burst blend 4.3/10; 10.5s, IT.
italy_18_365_13252208_13262752 · in -37.0 dBFS · gain +17.0 dB · eurospeech-01774
Contemplation ↓  /  Emotional Numbnessidentity +0.02 emotion 163 %   proxy_spearman__PXR__T0.40__C0.25__INTERNAL · #8

This chain comes from the proxy rule: the same two-sided test as above, but because Emotional Numbness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Emotional Numbness around average — 0.44, lower than 56 % of clips in this corpus — and ends with it strongly present at 0.84, higher than 84 % of clips in this corpus. That is a total rise of 0.40.

At the same time Contemplation goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.58 (higher than 58 % of clips in this corpus), a change of -0.41. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.18, then +0.16, then +0.06 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.85 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.85 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.85. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 47 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.647 before conversion and 0.666 after — it rose by 0.020. Neighbour-to-neighbour the worst pair went 0.719 → 0.674. (The earlier render, with segment 1 left raw, scores 0.594 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.401 in the original and +0.653 after conversion — 163 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Contemplation, -0.414 became -0.366.

Quality. Mean predicted overall quality across the segments went 2.76 → 3.02 (+0.26) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.647 → 0.666 +0.020identity cos neighbours 0.719 → 0.674d_b rescored +0.401 → +0.653d_a rescored -0.414 → -0.366d_a mined -0.413d_b mined 0.401min_cos_consec (site) 0.8469min_cos_anchor (site) 0.8546dataset emolialang enspeaker EN_0W0Q19H-OP8total 46.2schain gain +2.0 dBseam step 1.0 dBcrossfades 150/100/150 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a middle-aged masculine voice · neutral-toned, balanced body, measured, slightly relaxed
(contemplation, concentration, triumph · very low-energy, steady, frequent disfluency, monologue) To have idea meritocratic decision making. To try to come up with the best collective decisions. To, to bring the country together in an idea meritocratic way. (low mumble) Uhm, I, I think the period that we're in reminds me very much of 1937.
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly negative, neutral stance, slightly guarded; reads as contemplation, concentration, triumph; style: monologue, didactic; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 2.0/10; 20.9s, EN.
EN_0W0Q19H-OP8_W000003 · in -21.6 dBFS · gain +1.6 dB · emolia-01506
(fear · subdued, fairly steady, some disfluency, monologue) If I was to pick an analogous period of time, because at that time, it was after the financial crisis. We had 29 to 32, which is the equivalent of 2008 to 2009. They printed a lot of money, asset prices went up, interest rates went to zero.
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fear; style: monologue, formal; average recording, no background noise; genuineness 2.5/6; vocal-burst blend 0.9/10; 16.0s, EN.
EN_0W0Q19H-OP8_W000004 · in -20.3 dBFS · gain +0.3 dB · emolia-01506
(normally alert, fairly steady, little disfluency, monologue) And we had a large wealth gap which created populism. Analogous.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, casual; good recording, no background noise; genuineness 1.4/6; vocal-burst blend 1.2/10; 4.8s, EN.
EN_0W0Q19H-OP8_W000005 · in -20.0 dBFS · gain -0.0 dB · emolia-01506
(very low-energy, steady, little disfluency, monologue) The central bank begins to tighten monetary policy. And then we had
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, little disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; good recording, no background noise; genuineness 1.2/6; vocal-burst blend 2.4/10; 5.1s, EN.
EN_0W0Q19H-OP8_W000006 · in -21.6 dBFS · gain +1.6 dB · emolia-01506
Concentration ↓  /  Hope Enthusiasm Optimismidentity +0.43 emotion 115 %   proxy_spearman__PXR__T0.40__C0.25__INTERNAL · #9

This chain comes from the proxy rule: the same two-sided test as above, but because Hope Enthusiasm Optimism is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Hope Enthusiasm Optimism around average — 0.44, lower than 56 % of clips in this corpus — and ends with it strongly present at 0.88, higher than 88 % of clips in this corpus. That is a total rise of 0.44.

At the same time Concentration goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.58 (higher than 58 % of clips in this corpus), a change of -0.41. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.25, then -0.01, then +0.20 — not a clean run: step 2 moves back the other way by 0.01 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.26 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.24 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.26, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 41 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.250 before conversion and 0.679 after — it rose by 0.429. Neighbour-to-neighbour the worst pair went 0.268 → 0.736. (The earlier render, with segment 1 left raw, scores 0.576 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Hope Enthusiasm Optimism moved +0.443 in the original and +0.511 after conversion — 115 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Concentration, -0.406 became -0.432.

Quality. Mean predicted overall quality across the segments went 3.08 → 3.22 (+0.14) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.250 → 0.679 +0.429identity cos neighbours 0.268 → 0.736d_b rescored +0.443 → +0.511d_a rescored -0.406 → -0.432d_a mined -0.407d_b mined 0.442min_cos_consec (site) 0.2384min_cos_anchor (site) 0.2583dataset emolialang enspeaker EN_VKdqBiqFdWytotal 40.1schain gain +2.2 dBseam step 3.5 dBcrossfades 150/150/150 ms
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normally alert, slightly relaxed, light breath
(concentration · normal-paced, fairly steady, some disfluency, monologue) That number is then tied to that number on the right hand side. And so by scrubbing across the screen as they call it, you can get a better understanding of what's happening with the rest of the equation.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: monologue, formal; good recording, no background noise; genuineness 1.4/6; vocal-burst blend 0.6/10; 11.1s, EN.
EN_VKdqBiqFdWy_W000052 · in -19.4 dBFS · gain -0.7 dB · emolia-02424
(astonishment surprise · normal-paced, fairly steady, some disfluency, casual) Another, (low mumble) uh, something that he mentions with this that I didn't realize is an application called Solver by a company called Aqualia and it's in the Mac App Store right now and you can play with this same idea where, (low mumble) uhm,
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as astonishment surprise; style: casual, didactic; good recording, quiet background; genuineness 4.1/6; vocal-burst blend 3.7/10; 13.0s, EN.
EN_VKdqBiqFdWy_W000053 · in -18.2 dBFS · gain -1.8 dB · emolia-02424
(contentment · brisk, fairly steady, some disfluency, casual) Where you've got numbers, but it's different than a spreadsheet, so to speak. Let me try playing this introduction video. It's only 38 seconds long. You get a better idea of what this is. (low mumble) Uhm, it's wild stuff.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as contentment; style: casual, authoritative; good recording, quiet background; genuineness 2.7/6; vocal-burst blend 2.8/10; 10.0s, EN.
EN_VKdqBiqFdWy_W000054 · in -20.9 dBFS · gain +0.9 dB · emolia-02424
(normal-paced, steady, no disfluency, formal) Solver helps you do quick calculations and figure stuff out. Just type your problem and Solver shows you the answer.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: formal, whispered; average recording, no background noise; genuineness 0.9/6; vocal-burst blend 1.2/10; 6.7s, EN.
EN_VKdqBiqFdWy_W000055 · in -19.7 dBFS · gain -0.3 dB · emolia-02424
Interest ↓  /  Fatigue Exhaustionidentity +0.03 emotion 78 %   proxy_spearman__PXR__T0.40__C0.25__INTERNAL · #10

This chain comes from the proxy rule: the same two-sided test as above, but because Fatigue Exhaustion is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Fatigue Exhaustion essentially absent — 0.07, lower than 93 % of clips in this corpus — and ends with it strongly present at 0.85, higher than 85 % of clips in this corpus. That is a total rise of 0.78.

At the same time Interest goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.47 (lower than 53 % of clips in this corpus), a change of -0.51. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.17, then +0.18, then +0.22, then +0.20 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.82 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.84 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.82. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 61 s · fr · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.675 before conversion and 0.701 after — it rose by 0.026. Neighbour-to-neighbour the worst pair went 0.828 → 0.819. (The earlier render, with segment 1 left raw, scores 0.593 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fatigue Exhaustion moved +0.781 in the original and +0.607 after conversion — 78 % of the delta retained, which is most of it. On the other named axis, Interest, -0.507 became -0.515.

Quality. Mean predicted overall quality across the segments went 2.95 → 3.12 (+0.17) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.675 → 0.701 +0.026identity cos neighbours 0.828 → 0.819d_b rescored +0.781 → +0.607d_a rescored -0.507 → -0.515d_a mined -0.507d_b mined 0.781min_cos_consec (site) 0.8441min_cos_anchor (site) 0.8170dataset emolialang frspeaker FR_6jHVwjh7usktotal 59.9schain gain +0.4 dBseam step 1.4 dBcrossfades 150/100/150/150 ms
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: an elderly masculine voice · neutral-toned, fairly smooth, balanced body, average recording, quiet background, moderate pitch range
(interest, sourness, pain · normal-paced, normally alert, slightly relaxed, monologue) En Angleterre, Givens, Stanley Givens, qui était un économiste, a réalisé une machine qui utilisait l'algèbre de boules pour faire des raisonnements logiques avec cette idée de reproduire la pensée. Mais on a quand même une (low mumble) (low mumble) personnalité assez singulière qui a joué un rôle très important en informatique et qui a été un peu considéré comme le précurseur de la discipline. C'est Alan Turing avec les articles qu'il a écrits en euh, (low mumble) (low mumble) voilà. Est-ce que, qu'est-ce que ça voudrait dire que (low mumble) les
full caption & clip details
An elderly masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as interest, sourness, pain; style: monologue, storytelling; average recording, quiet background; genuineness 2.8/6; vocal-burst blend 4.1/10; 30.0s, FR.
FR_6jHVwjh7usk_W000064 · in -17.1 dBFS · gain -2.9 dB · emolia-02679
(confusion, contemplation · normal-paced, normally alert, neutral tension, conversational) Et je crois que c'est important, c'est de bien comprendre, et ça rejoindra peut-être ce que dit (low mumble) Jean-François ici, qu'Alain Turing (ahem) avait cette idée,
full caption & clip details
An elderly masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; reads as confusion, contemplation; style: conversational, casual; average recording, quiet background; genuineness 4.4/6; vocal-burst blend 4.6/10; 10.5s, FR.
FR_6jHVwjh7usk_W000065 · in -18.4 dBFS · gain -1.6 dB · emolia-02679
(anger · measured, normally alert, slightly relaxed, storytelling) que, effectivement, on devait pouvoir fabriquer des machines qui reproduisent la pensée.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, audible breath; affect is neutral, slightly dominant, neutral openness; reads as anger; style: storytelling, conversational; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 3.8/10; 6.2s, FR.
FR_6jHVwjh7usk_W000066 · in -16.5 dBFS · gain -3.5 dB · emolia-02679
(confusion, disappointment · measured, normally alert, slightly relaxed, didactic) Et il s'est demandé comment faire. Et là, il a envisagé tout un tas de possibilités. Et très tôt, il s'est dit pour que l'on simule la pensée.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; very clear, frequent disfluency, moderate pitch range, audible breath; affect is neutral, slightly dominant, slightly guarded; reads as confusion, disappointment; style: didactic, dramatic; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 3.0/10; 8.6s, FR.
FR_6jHVwjh7usk_W000067 · in -17.5 dBFS · gain -2.5 dB · emolia-02679
(brisk, energised, slightly relaxed, storytelling) On va essayer de prendre des tâches qui sont faciles à simuler. C'est-à-dire qu'il a essayé de réduire l'ambition.
full caption & clip details
An adult masculine voice; delivery is energised, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; very clear, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, neutral openness; no dominant emotion; style: storytelling, dramatic; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 2.3/10; 5.3s, FR.
FR_6jHVwjh7usk_W000068 · in -16.4 dBFS · gain -3.6 dB · emolia-02679
Longing ↓  /  Hope Enthusiasm Optimismidentity +0.03 emotion 66 %   proxy_spearman__PXR__T0.40__C0.25__INTERNAL · #11

This chain comes from the proxy rule: the same two-sided test as above, but because Hope Enthusiasm Optimism is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Hope Enthusiasm Optimism around average — 0.57, higher than 57 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.42.

At the same time Longing goes the other way, from 0.95 (higher than 95 % of clips in this corpus) to 0.54 (higher than 54 % of clips in this corpus), a change of -0.41. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.21, then +0.11, then +0.09, then +0.01 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.72 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.74 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.72, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 56 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.647 before conversion and 0.681 after — it rose by 0.034. Neighbour-to-neighbour the worst pair went 0.678 → 0.696. (The earlier render, with segment 1 left raw, scores 0.595 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Hope Enthusiasm Optimism moved +0.418 in the original and +0.276 after conversion — 66 % of the delta retained. On the other named axis, Longing, -0.406 became -0.456.

Quality. Mean predicted overall quality across the segments went 2.64 → 2.99 (+0.35) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.647 → 0.681 +0.034identity cos neighbours 0.678 → 0.696d_b rescored +0.418 → +0.276d_a rescored -0.406 → -0.456d_a mined -0.406d_b mined 0.419min_cos_consec (site) 0.7414min_cos_anchor (site) 0.7178dataset emolialang enspeaker EN_B00039_S03752total 54.5schain gain +3.5 dBseam step 0.7 dBcrossfades 150/150/150/100 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, balanced body
(longing, affection · measured, normally alert, slightly relaxed, casual) Boy. From his family, all his friends. And then putting him with another family, he's like, nope. And then...
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, frequent disfluency, moderate pitch range, normal breath; affect is neutral, slightly dominant, neutral openness; reads as longing, affection; style: casual, monologue; good recording, quiet background; genuineness 3.1/6; vocal-burst blend 1.2/10; 6.3s, EN.
EN_B00039_S03752_W000020 · in -21.0 dBFS · gain +1.0 dB · emolia-00995
(infatuation, astonishment surprise, pleasure ecstasy · brisk, normally alert, neutral tension, casual) Here he is, so I think it's weird for him. This is the first time he's ever been like inside of a house before he was an outside dog.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as infatuation, astonishment surprise, pleasure ecstasy; style: casual, conversational; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 6.5/10; 6.1s, EN.
EN_B00039_S03752_W000021 · in -19.1 dBFS · gain -0.9 dB · emolia-00995
(contentment, jealousy and envy, pleasure ecstasy · normal-paced, energised, neutral tension, casual) He's only about like 10 weeks old, so. There's also new to him, like, his brothers and sisters that he had, and his mom and dad that he grew up with, like, in the outside. He doesn't have them anymore. Now it's just Lee and I. So, I think he's gonna do alright. He's pretty happy right now. Check him out. Look, look at this. He's just playing with his toys.
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as contentment, jealousy and envy, pleasure ecstasy; style: casual, conversational; average recording, quiet background; genuineness 5.0/6; vocal-burst blend 10.0/10; 16.9s, EN.
EN_B00039_S03752_W000022 · in -19.3 dBFS · gain -0.7 dB · emolia-00995
(elation, hope enthusiasm optimism, pleasure ecstasy · normal-paced, very low-energy, neutral tension, casual) Oh, we, we pooped him out. Yeah, good. But we're gonna set up that cage. He has a little bed here. We're gonna make the bed more puppy friendly. Or the room more puppy friendly. Yes. Maybe there's food and stuff. I have all these boxes I gotta unbox. I got my computer over there. We gotta organize his room. Plus, Leah.
full caption & clip details
A young adult masculine voice; delivery is very low-energy, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as elation, hope enthusiasm optimism, pleasure ecstasy; style: casual, conversational; below-average recording, some background noise; mildly explicit content; genuineness 4.7/6; vocal-burst blend 6.8/10; 16.5s, EN.
EN_B00039_S03752_W000023 · in -18.6 dBFS · gain -1.4 dB · emolia-00995
(hope enthusiasm optimism, elation, pleasure ecstasy · normal-paced, normally alert, relaxed, casual) We're getting a new room. We're having a new room. So more information to come with that. But let's go ahead and close out this vlog with some comments you guys submitted on yesterday's vlog. I'm excited to showcase some of the
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, relaxed, fairly steady; timbre is neutral-toned, slightly bright, slightly rough, balanced body; somewhat unclear, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as hope enthusiasm optimism, elation, pleasure ecstasy; style: casual, conversational; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 8.2/10; 9.5s, EN.
EN_B00039_S03752_W000024 · in -18.3 dBFS · gain -1.7 dB · emolia-00995
Pain ↓  /  Concentrationidentity +0.22 emotion 93 %   proxy_spearman__PXR__T0.40__C0.25__INTERNAL · #12

This chain comes from the proxy rule: the same two-sided test as above, but because Concentration is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Concentration below average — 0.41, lower than 59 % of clips in this corpus — and ends with it strongly present at 0.89, higher than 89 % of clips in this corpus. That is a total rise of 0.48.

At the same time Pain goes the other way, from 0.86 (higher than 86 % of clips in this corpus) to 0.35 (lower than 65 % of clips in this corpus), a change of -0.51. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.14, then +0.22, then +0.13 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.28 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.20 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.28, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 38 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.276 before conversion and 0.497 after — it rose by 0.221. Neighbour-to-neighbour the worst pair went 0.175 → 0.620. (The earlier render, with segment 1 left raw, scores 0.542 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.479 in the original and +0.446 after conversion — 93 % of the delta retained, which is essentially all of it. On the other named axis, Pain, -0.514 became +0.348.

Quality. Mean predicted overall quality across the segments went 2.60 → 2.96 (+0.37) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.276 → 0.497 +0.221identity cos neighbours 0.175 → 0.620d_b rescored +0.479 → +0.446d_a rescored -0.514 → +0.348d_a mined -0.514d_b mined 0.480min_cos_consec (site) 0.1952min_cos_anchor (site) 0.2777dataset emolialang enspeaker EN_dxskzgGNzYMtotal 37.2schain gain +2.6 dBseam step 3.4 dBcrossfades 150/100/150 ms
Script — 4 chunks, 3 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · fairly smooth, quiet background, fairly steady
(normal-paced, normally alert, slightly relaxed, monologue) Money coaching partnerships that the library has with the two different organizations. And if you just go to our events calendar, they're listed five days a week.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, fairly smooth, thin; somewhat unclear, some disfluency, moderate pitch range, normal breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: monologue; below-average recording, quiet background; genuineness 2.7/6; vocal-burst blend 1.5/10; 10.4s, EN.
EN_dxskzgGNzYM_W000175 · in -17.9 dBFS · gain -2.1 dB · emolia-01103
(slow, very low-energy, relaxed, casual) And you click on, on the link, the registration link, and it takes you to their websites for (ahem) setting up meetings. So, (low mumble) let me give you those links too.
full caption & clip details
A middle-aged somewhat feminine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is slightly cool, slightly dark, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly positive, slightly submissive, neutral openness; no dominant emotion; style: casual, whispered; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 0.0/10; 12.4s, EN.
EN_dxskzgGNzYM_W000176 · in -17.3 dBFS · gain -2.7 dB · emolia-01103
(confusion, thankfulness gratitude, doubt · normal-paced, normally alert, slightly relaxed, conversational) (ahem) So someone just asked a question about how much of your savings rate should go towards investments. So it really, it's one of those, it depends. (low mumble)
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as confusion, thankfulness gratitude, doubt; style: conversational, casual; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 2.6/10; 7.1s, EN.
EN_dxskzgGNzYM_W000177 · in -16.7 dBFS · gain -3.3 dB · emolia-01103
(normal-paced, normally alert, slightly relaxed, conversational) So you want to look at your whole take home pay 20% going towards, (low mumble) um, savings. It depends on the situation. So.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; no dominant emotion; style: conversational, casual; good recording, quiet background; genuineness 2.0/6; vocal-burst blend 0.9/10; 7.8s, EN.
EN_dxskzgGNzYM_W000178 · in -19.0 dBFS · gain -1.0 dB · emolia-01103
Fatigue Exhaustion ↓  /  Malevolence Maliceidentity +0.00 emotion 55 %   proxy_spearman__PXR__T0.40__C0.25__INTERNAL · #13

This chain comes from the proxy rule: the same two-sided test as above, but because Malevolence Malice is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Malevolence Malice below average — 0.30, lower than 70 % of clips in this corpus — and ends with it strongly present at 0.86, higher than 86 % of clips in this corpus. That is a total rise of 0.56.

At the same time Fatigue Exhaustion goes the other way, from 0.85 (higher than 85 % of clips in this corpus) to 0.40 (lower than 60 % of clips in this corpus), a change of -0.45. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.21, then +0.22, then +0.13 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.83 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.86 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.83. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 31 s · zh · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.837 before conversion and 0.837 after — it rose by 0.001. Neighbour-to-neighbour the worst pair went 0.842 → 0.805. (The earlier render, with segment 1 left raw, scores 0.793 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Malevolence Malice moved +0.559 in the original and +0.309 after conversion — 55 % of the delta retained. On the other named axis, Fatigue Exhaustion, -0.449 became +0.290.

Quality. Mean predicted overall quality across the segments went 3.03 → 3.16 (+0.13) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.837 → 0.837 +0.001identity cos neighbours 0.842 → 0.805d_b rescored +0.559 → +0.309d_a rescored -0.449 → +0.290d_a mined -0.449d_b mined 0.560min_cos_consec (site) 0.8558min_cos_anchor (site) 0.8317dataset emolialang zhspeaker ZH_B00000_S07812total 29.9schain gain +1.7 dBseam step 1.4 dBcrossfades 100/100/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, average recording, no background noise, measured, normally alert, slightly relaxed
(no disfluency, somewhat unclear, narration, monologue) 头脑极度开放的人知道呢,找到问题的所有答案很重要。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, monologue; average recording, no background noise; genuineness 1.7/6; vocal-burst blend 3.4/10; 5.1s, ZH.
ZH_B00000_S07812_W000020 · in -19.3 dBFS · gain -0.7 dB · emolia-03267
(no disfluency, clear, narration, monologue) 但是呢提出正确的问题,并且向相关领域的专业人士请教也很重要。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, monologue; average recording, no background noise; genuineness 1.2/6; vocal-burst blend 3.1/10; 6.7s, ZH.
ZH_B00000_S07812_W000021 · in -18.6 dBFS · gain -1.4 dB · emolia-03267
(no disfluency, clear, monologue, didactic) 我们愿意请教,就说明我们知道自己还是有一些东西是不知道的,而且呢我们不会因为自己不知道某些事情而感到好像不好意思。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, didactic; average recording, no background noise; genuineness 1.9/6; vocal-burst blend 4.4/10; 10.9s, ZH.
ZH_B00000_S07812_W000022 · in -19.2 dBFS · gain -0.8 dB · emolia-03267
(some disfluency, somewhat unclear, monologue, authoritative) 这样呢我们是不是就能做出更好的决策了,因为跟任何人知道的任何东西相比。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, authoritative; average recording, no background noise; genuineness 3.0/6; vocal-burst blend 4.7/10; 7.7s, ZH.
ZH_B00000_S07812_W000023 · in -18.1 dBFS · gain -1.9 dB · emolia-03267
Intoxication Altered States of Consciousness ↓  /  Malevolence Maliceidentity −0.06 emotion REVERSED   proxy_spearman__PXR__T0.40__C0.25__INTERNAL · #14

This chain comes from the proxy rule: the same two-sided test as above, but because Malevolence Malice is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Malevolence Malice below average — 0.40, lower than 60 % of clips in this corpus — and ends with it strongly present at 0.82, higher than 82 % of clips in this corpus. That is a total rise of 0.42.

At the same time Intoxication Altered States of Consciousness goes the other way, from 0.60 (higher than 60 % of clips in this corpus) to 0.19 (lower than 81 % of clips in this corpus), a change of -0.41. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.21, then +0.21 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.89 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.91 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.89. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 14 s · zh · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.899 before conversion and 0.836 after — it fell by 0.063. Neighbour-to-neighbour the worst pair went 0.913 → 0.855. (The earlier render, with segment 1 left raw, scores 0.836 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

The emotional move did not survive. Re-scored end to end, Malevolence Malice moved +0.422 in the original and -0.316 after conversion — it changed direction. On this chain the corrected audio is not an improvement. On the other named axis, Intoxication Altered States of Consciousness, -0.414 became -0.551.

Quality. Mean predicted overall quality across the segments went 2.88 → 2.90 (+0.02) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.899 → 0.836 -0.063identity cos neighbours 0.913 → 0.855d_b rescored +0.422 → -0.316d_a rescored -0.414 → -0.551d_a mined -0.414d_b mined 0.422min_cos_consec (site) 0.9117min_cos_anchor (site) 0.8921dataset emolialang zhspeaker ZH_B00029_S09866total 13.2schain gain +1.7 dBseam step 0.7 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, no background noise, measured, normally alert, slightly relaxed, no disfluency
(fairly steady, formal, narration) 但归根结底,这抗拒之中,包含着诱惑。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, narration; very good recording, no background noise; genuineness 0.2/6; vocal-burst blend 5.0/10; 3.5s, ZH.
ZH_B00029_S09866_W000133 · in -21.3 dBFS · gain +1.3 dB · emolia-03568
(fairly steady, narration, formal) 而恰恰是这种让人捉摸不定的态度,刺激了他的欲望。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, formal; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 4.7/10; 4.3s, ZH.
ZH_B00029_S09866_W000134 · in -22.5 dBFS · gain +2.5 dB · emolia-03568
(steady, formal, narration) 无论如何,他已经有了伙伴,这出戏可以演出了神速的友谊。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, narration; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 3.7/10; 5.9s, ZH.
ZH_B00029_S09866_W000135 · in -21.8 dBFS · gain +1.8 dB · emolia-03568
Distress ↓  /  Contentmentidentity +0.01 emotion 75 %   proxy_spearman__PXR__T0.40__C0.25__INTERNAL · #15

This chain comes from the proxy rule: the same two-sided test as above, but because Contentment is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Contentment around average — 0.52, higher than 52 % of clips in this corpus — and ends with it at the very top of the corpus at 0.92, higher than 92 % of clips in this corpus. That is a total rise of 0.41.

At the same time Distress goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.44 (lower than 56 % of clips in this corpus), a change of -0.54. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.25, then +0.05, then +0.08, then +0.03 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the eurospeech clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 66 s · it · eurospeech

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.882 before conversion and 0.891 after — it rose by 0.009. Neighbour-to-neighbour the worst pair went 0.900 → 0.908. (The earlier render, with segment 1 left raw, scores 0.781 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contentment moved +0.409 in the original and +0.307 after conversion — 75 % of the delta retained, which is most of it. On the other named axis, Distress, -0.544 became -0.530.

Quality. Mean predicted overall quality across the segments went 2.89 → 3.26 (+0.37) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.882 → 0.891 +0.009identity cos neighbours 0.900 → 0.908d_b rescored +0.409 → +0.307d_a rescored -0.544 → -0.530d_a mined -0.544d_b mined 0.409min_cos_consec (site) —min_cos_anchor (site) —dataset eurospeechlang itspeaker italy_18_65total 64.9schain gain +1.9 dBseam step 2.6 dBcrossfades 100/100/100/100 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · slightly cool, slightly bright, fairly smooth, thin, quiet background, brisk, moderately variable, wide pitch range
(distress, disappointment, anger · energised, neutral tension, some disfluency, dramatic) che analizza la situazione italiana in materia di contrasto alla violenza sulle donne, ha evidenziato che nel nostro Paese la normativa esistente in concreto non viene applicata e che ancora, malgrado le risorse economiche stanziate e le azioni messe in campo per affrontare il fenomeno,
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; clear, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, fairly guarded; reads as distress, disappointment, anger; style: dramatic, ranting; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 5.4/10; 18.3s, IT.
italy_18_65_7023424_7041712 · in -11.4 dBFS · gain -8.6 dB · eurospeech-01786
(shame, disappointment, helplessness · energised, neutral tension, some disfluency, dramatic) la distribuzione e l'applicazione disomogenea delle stesse risorse sul territorio nazionale fa sì che la tutela dei diritti delle vittime di violenza non sia effettiva.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as shame, disappointment, helplessness; style: dramatic, casual; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 5.6/10; 11.8s, IT.
italy_18_65_7041712_7053472 · in -12.1 dBFS · gain -7.9 dB · eurospeech-01786
(relief, affection, shame · normally alert, slightly relaxed, some disfluency, dramatic) Con questa mozione oggi vogliamo rendere possibile che quanto stanziato annualmente per i centri antiviolenza e le strutture di accoglienza per mamme e minori
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as relief, affection, shame; style: dramatic, cartoonish; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 4.4/10; 10.4s, IT.
italy_18_65_7053472_7063920 · in -11.7 dBFS · gain -8.3 dB · eurospeech-01786
(shame, concentration, pride · energised, neutral tension, some disfluency, dramatic) sia erogato regolarmente, senza ritardi e vincolato all'assunzione di impegni precisi, all'individuazione delle priorità e alla valutazione dei risultati ottenuti.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; clear, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, guarded; reads as shame, concentration, pride; style: dramatic, casual; below-average recording, quiet background; genuineness 3.2/6; vocal-burst blend 7.4/10; 11.5s, IT.
italy_18_65_7063920_7075456 · in -10.9 dBFS · gain -9.1 dB · eurospeech-01786
(contentment, infatuation, thankfulness gratitude · energised, neutral tension, almost no disfluency, dramatic) È necessario colmare le lacune di applicazione delle leggi in materia, introducendo modalità di attuazione volte a snellire i procedimenti con chiara individuazione della norma ed imposizione della stessa.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; clear, almost no disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, guarded; reads as contentment, infatuation, thankfulness gratitude; style: dramatic, cartoonish; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 6.4/10; 13.4s, IT.
italy_18_65_7075456_7088816 · in -10.6 dBFS · gain -9.4 dB · eurospeech-01786
Emotional Numbness ↓  /  Painidentity −0.10 emotion 38 %   proxy_spearman__PXR__T0.40__C0.25__INTERNAL · #16

This chain comes from the proxy rule: the same two-sided test as above, but because Pain is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Pain barely there — 0.20, lower than 80 % of clips in this corpus — and ends with it clearly present at 0.72, higher than 72 % of clips in this corpus. That is a total rise of 0.52.

At the same time Emotional Numbness goes the other way, from 0.95 (higher than 95 % of clips in this corpus) to 0.52 (higher than 52 % of clips in this corpus), a change of -0.43. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.15, then +0.00, then +0.20, then +0.17 — a plateau around step 2, where it barely moves.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.93 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.96 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.93), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 28 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.884 before conversion and 0.785 after — it fell by 0.098. Neighbour-to-neighbour the worst pair went 0.875 → 0.748. (The earlier render, with segment 1 left raw, scores 0.511 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Pain moved +0.520 in the original and +0.196 after conversion — 38 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Emotional Numbness, -0.426 became -0.379.

Quality. Mean predicted overall quality across the segments went 2.75 → 2.91 (+0.17) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.884 → 0.785 -0.098identity cos neighbours 0.875 → 0.748d_b rescored +0.520 → +0.196d_a rescored -0.426 → -0.379d_a mined -0.426d_b mined 0.520min_cos_consec (site) 0.9622min_cos_anchor (site) 0.9294dataset emolialang enspeaker EN_X_Mq4PvxgOototal 26.4schain gain +1.4 dBseam step 1.3 dBcrossfades 100/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(emotional numbness, contempt, disgust · formal, monologue) To them the sense of subject-object perception was illusory and a sign of ignorance
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, contempt, disgust; style: formal, monologue; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 1.2/10; 4.6s, EN.
EN_X_Mq4PvxgOo_W000141 · in -14.7 dBFS · gain -5.3 dB · emolia-02621
(emotional numbness, doubt · formal, authoritative) However, the individual's sense of self was not a complete illusion since it was derived from the universal beingness that is Brahman
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, doubt; style: formal, authoritative; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.0/10; 7.0s, EN.
EN_X_Mq4PvxgOo_W000142 · in -14.2 dBFS · gain -5.8 dB · emolia-02621
(infatuation · formal, monologue) Ramanuja sa Vishnu as a personification of Brahman
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as infatuation; style: formal, monologue; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 1.5/10; 3.3s, EN.
EN_X_Mq4PvxgOo_W000143 · in -14.3 dBFS · gain -5.7 dB · emolia-02621
(formal, authoritative) Dvaita refers to a theistic sub-school in Vedanta tradition of Hindu philosophy
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.2/10; 4.9s, EN.
EN_X_Mq4PvxgOo_W000145 · in -13.9 dBFS · gain -6.1 dB · emolia-02621
(formal, authoritative) Also called as Tattvavada and Bimbapratibimbavada, the Dvaita sub-school was founded by the 13th-century scholar Madhvacharya
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.0/10; 7.4s, EN.
EN_X_Mq4PvxgOo_W000146 · in -15.5 dBFS · gain -4.5 dB · emolia-02621
Pain ↓  /  Angeridentity +0.32 emotion 70 %   proxy_spearman__PXR__T0.40__C0.25__INTERNAL · #17

This chain comes from the proxy rule: the same two-sided test as above, but because Anger is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Anger around average — 0.48, lower than 52 % of clips in this corpus — and ends with it strongly present at 0.89, higher than 89 % of clips in this corpus. That is a total rise of 0.41.

At the same time Pain goes the other way, from 0.85 (higher than 85 % of clips in this corpus) to 0.35 (lower than 65 % of clips in this corpus), a change of -0.50. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.23, then +0.01, then +0.18 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores -0.04 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.06 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (-0.04, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 21 s · snippets

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.045 before conversion and 0.369 after — it rose by 0.324. Neighbour-to-neighbour the worst pair went 0.045 → 0.369. (The earlier render, with segment 1 left raw, scores 0.276 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Anger moved +0.414 in the original and +0.292 after conversion — 70 % of the delta retained, which is most of it. On the other named axis, Pain, -0.502 became -0.542.

Quality. Mean predicted overall quality across the segments went 2.79 → 2.87 (+0.08) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.045 → 0.369 +0.324identity cos neighbours 0.045 → 0.369d_b rescored +0.414 → +0.292d_a rescored -0.502 → -0.542d_a mined -0.502d_b mined 0.415min_cos_consec (site) 0.0587min_cos_anchor (site) -0.0358dataset snippetslang undspeaker batch276_part0_batch276_patotal 20.1schain gain +2.3 dBseam step 4.6 dBcrossfades 150/150/100 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, light breath
(normal-paced, normally alert, slightly relaxed, storytelling) General Gamelan's French forces mass on the Belgian frontier.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; no dominant emotion; style: storytelling, narration; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 1.2/10; 3.5s.
batch276_part0_batch276_part0_chunk_950_1_830576 · in -25.2 dBFS · gain +5.2 dB · snippets-00917
(awe, fear · brisk, normally alert, neutral tension, storytelling) fact that this mighty huge red army suffered humiliations on the battlefield in these forests and snows.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, full; average clarity, almost no disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as awe, fear; style: storytelling, dramatic; good recording, no background noise; genuineness 1.9/6; vocal-burst blend 1.4/10; 7.2s.
batch276_part0_batch276_part0_chunk_950_1_830654 · in -25.1 dBFS · gain +5.1 dB · snippets-00917
(elation, triumph, hope enthusiasm optimism · normal-paced, energised, slightly tense, casual) probably single figure that best embodies the philosophy of moderation in Montesquieu even though he was not as well read in Montesquieu as these other
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, slightly tense, moderately variable; timbre is neutral-toned, slightly bright, slightly rough, thin; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as elation, triumph, hope enthusiasm optimism; style: casual, storytelling; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 4.3/10; 3.6s.
batch276_part0_batch276_part0_chunk_950_1_830719 · in -24.3 dBFS · gain +4.3 dB · snippets-00917
(normal-paced, normally alert, slightly relaxed, conversational) means lightning war that's the translation and that's exactly what this was tanks and motorized infantry
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; no dominant emotion; style: conversational, casual; good recording, no background noise; genuineness 2.7/6; vocal-burst blend 3.1/10; 6.3s.
batch276_part0_batch276_part0_chunk_950_1_830745 · in -26.6 dBFS · gain +6.6 dB · snippets-00917
Contemplation ↓  /  Contentmentidentity +0.13 emotion 69 %   proxy_spearman__PXR__T0.40__C0.25__INTERNAL · #18

This chain comes from the proxy rule: the same two-sided test as above, but because Contentment is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Contentment around average — 0.54, higher than 54 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.44.

At the same time Contemplation goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.54 (higher than 54 % of clips in this corpus), a change of -0.44. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.25, then +0.17, then +0.02 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.26 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.37 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.26, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 44 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.292 before conversion and 0.422 after — it rose by 0.130. Neighbour-to-neighbour the worst pair went 0.396 → 0.481. (The earlier render, with segment 1 left raw, scores 0.335 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contentment moved +0.438 in the original and +0.303 after conversion — 69 % of the delta retained. On the other named axis, Contemplation, -0.443 became -0.365.

Quality. Mean predicted overall quality across the segments went 2.72 → 2.97 (+0.25) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.292 → 0.422 +0.130identity cos neighbours 0.396 → 0.481d_b rescored +0.438 → +0.303d_a rescored -0.443 → -0.365d_a mined -0.443d_b mined 0.438min_cos_consec (site) 0.3660min_cos_anchor (site) 0.2632dataset emolialang enspeaker EN_tA2xfSqaPE0total 42.6schain gain +4.3 dBseam step 1.7 dBcrossfades 100/150/150 ms
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, fairly steady, moderate pitch range, light breath
(contemplation, doubt · normal-paced, normally alert, slightly relaxed, casual) We'll come to know that there's not these, (ahem) we may be in different places on the planet.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as contemplation, doubt; style: casual, conversational; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 2.3/10; 6.2s, EN.
EN_tA2xfSqaPE0_W000325 · in -18.0 dBFS · gain -2.0 dB · emolia-01890
(emotional numbness · normal-paced, normally alert, slightly relaxed, casual) And supposedly in different time zones, but it's really just one time zone. And that's time zone art, so.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as emotional numbness; style: casual, monologue; good recording, quiet background; genuineness 3.2/6; vocal-burst blend 3.0/10; 6.9s, EN.
EN_tA2xfSqaPE0_W000326 · in -18.3 dBFS · gain -1.7 dB · emolia-01890
(hope enthusiasm optimism, interest, contentment · normal-paced, normally alert, neutral tension, casual) Times on Art and I wanted to say yeaah, that's where we meet actually, in the Times on Art. Yeaah. And I also want to say, you know, hello to Noam and Rafika, they are, you know,
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as hope enthusiasm optimism, interest, contentment; style: casual, conversational; average recording, quiet background; genuineness 5.4/6; vocal-burst blend 8.9/10; 11.6s, EN.
EN_tA2xfSqaPE0_W000328 · in -21.5 dBFS · gain +1.5 dB · emolia-01890
(contentment, relief, pleasure ecstasy · measured, very low-energy, neutral tension, casual) In, in, in the east side of the world, and, uh, (ahem) it's like 3 a.m. in Israel now, and it's 6 a.m. in Taiwan now, so it was harder for them to, to reach us, but they are here also, and with us. And, (low mumble) uh, yeah, and, (low mumble) uh, also wanted
full caption & clip details
A young adult masculine voice; delivery is very low-energy, measured, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as contentment, relief, pleasure ecstasy; style: casual, monologue; average recording, quiet background; genuineness 4.9/6; vocal-burst blend 8.6/10; 18.5s, EN.
EN_tA2xfSqaPE0_W000329 · in -20.5 dBFS · gain +0.5 dB · emolia-01890
Emotional Numbness ↓  /  Shameidentity +0.01 emotion 90 %   proxy_spearman__PXR__T0.40__C0.25__INTERNAL · #19

This chain comes from the proxy rule: the same two-sided test as above, but because Shame is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Shame around average — 0.47, lower than 53 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.52.

At the same time Emotional Numbness goes the other way, from 0.95 (higher than 95 % of clips in this corpus) to 0.44 (lower than 56 % of clips in this corpus), a change of -0.50. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.23, then +0.09, then +0.20 — an uneven climb, but always in the same direction.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.80 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.85 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.80. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 60 s · en · podcast

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.602 before conversion and 0.615 after — it rose by 0.014. Neighbour-to-neighbour the worst pair went 0.663 → 0.699. (The earlier render, with segment 1 left raw, scores 0.500 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Shame moved +0.521 in the original and +0.470 after conversion — 90 % of the delta retained, which is essentially all of it. On the other named axis, Emotional Numbness, -0.500 became -0.634.

Quality. Mean predicted overall quality across the segments went 2.60 → 2.99 (+0.40) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.602 → 0.615 +0.014identity cos neighbours 0.663 → 0.699d_b rescored +0.521 → +0.470d_a rescored -0.500 → -0.634d_a mined -0.502d_b mined 0.521min_cos_consec (site) 0.8494min_cos_anchor (site) 0.8008dataset podcastlang enspeaker 135755total 58.5schain gain +1.3 dBseam step 1.5 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-bright
(emotional numbness · measured, normally alert, slightly relaxed, casual) Because he was entering into his ministry. He had turned 30.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as emotional numbness; style: casual, monologue; average recording, quiet background; genuineness 2.3/6; vocal-burst blend 1.8/10; 3.7s, EN.
135755_00244912 · in -26.8 dBFS · gain +6.8 dB · podcast-01274
(measured, very low-energy, neutral tension, casual) And he didn't turn to 53, but he had them pull out Isaiah 61, what we call right.
full caption & clip details
An adult masculine voice; delivery is very low-energy, measured, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, frequent disfluency, wide pitch range, normal breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 1.7/10; 8.2s, EN.
135755_00245740 · in -26.2 dBFS · gain +6.2 dB · podcast-03524
(bitterness, anger, impatience and irritability · measured, highly aroused, neutral tension, dramatic) They didn't have numbers and verses and things. That's just for our reference. But he called for the scroll and he turned to this and he said, The Spirit of the Lord God is upon me, because the Lord hath anointed me to preach good tidings to the meek. He has sent me to bind up the brokenhearted. Are you brokenhearted
full caption & clip details
A middle-aged masculine voice; delivery is highly aroused, measured, neutral tension, variable; timbre is slightly cool, neutral-bright, rough, balanced body; average clarity, frequent disfluency, wide pitch range, normal breath; affect is negative, dominant, fairly guarded; reads as bitterness, anger, impatience and irritability; style: dramatic, authoritative; below-average recording, some background noise; genuineness 2.0/6; vocal-burst blend 0.6/10; 23.6s, EN.
135755_00246564 · in -21.5 dBFS · gain +1.5 dB · podcast-03535
(shame, contempt, contentment · brisk, highly aroused, slightly tense, dramatic) proclaim liberty to the captives. Do you feel captive to your own dumb decisions, your own willfulness? Do you feel like you are in a prison of your own making? Well, he's here to set you at liberty. The opening of prison to them that are bound, to proclaim the acceptable year of the Lord, the day of vengeance of our God, to comfort all that mourn.
full caption & clip details
A middle-aged masculine voice; delivery is highly aroused, brisk, slightly tense, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, full; very clear, some disfluency, wide pitch range, normal breath; affect is positive, dominant, fairly guarded; reads as shame, contempt, contentment; style: dramatic, authoritative; below-average recording, some background noise; genuineness 1.8/6; vocal-burst blend 2.8/10; 23.6s, EN.
135755_00249304 · in -19.1 dBFS · gain -0.9 dB · podcast-03534
Hope Enthusiasm Optimism ↓  /  Embarrassmentidentity −0.04 emotion 86 %   proxy_spearman__PXR__T0.40__C0.25__INTERNAL · #20

This chain comes from the proxy rule: the same two-sided test as above, but because Embarrassment is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Embarrassment around average — 0.56, higher than 56 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.43.

At the same time Hope Enthusiasm Optimism goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.51 (right about the corpus median), a change of -0.47. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.24, then -0.02, then +0.22 — not a clean run: step 2 moves back the other way by 0.02 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.90 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.91 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.90. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 82 s · en · podcast

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.853 before conversion and 0.813 after — it fell by 0.040. Neighbour-to-neighbour the worst pair went 0.893 → 0.813. (The earlier render, with segment 1 left raw, scores 0.703 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Embarrassment moved +0.433 in the original and +0.371 after conversion — 86 % of the delta retained, which is most of it. On the other named axis, Hope Enthusiasm Optimism, -0.473 became -0.511.

Quality. Mean predicted overall quality across the segments went 2.64 → 3.20 (+0.56) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.853 → 0.813 -0.040identity cos neighbours 0.893 → 0.813d_b rescored +0.433 → +0.371d_a rescored -0.473 → -0.511d_a mined -0.471d_b mined 0.433min_cos_consec (site) 0.9080min_cos_anchor (site) 0.8990dataset podcastlang enspeaker 903763total 81.3schain gain +4.5 dBseam step 1.0 dBcrossfades 150/150/150 ms
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, balanced body, average recording, quiet background, neutral tension, some disfluency, somewhat unclear
(hope enthusiasm optimism, malevolence malice, sexual lust · measured, energised, moderately variable, casual) And if you're willing to change your life, go and do whatever it takes to get to your daughter. But if you ever go back and sell one gram or one ball, you're just nothing but a snitch. (yawn) So
full caption & clip details
An adult masculine voice; delivery is energised, measured, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, wide pitch range, normal breath; affect is mildly negative, slightly dominant, fairly guarded; reads as hope enthusiasm optimism, malevolence malice, sexual lust; style: casual, storytelling; average recording, quiet background; mildly explicit content; genuineness 4.0/6; vocal-burst blend 2.9/10; 13.0s, EN.
903763_00199568 · in -19.5 dBFS · gain -0.5 dB · podcast-03708
(affection, elation, longing · normal-paced, normally alert, fairly steady, casual) I get home and I get a phone call from remember my friend Sandy Perkoff, the yacht broker. Well, we had he had a friend named Dave Jackson, who was also a yacht broker. But Dave was a really big yacht broker. He sold to Jimmy Buffett. He sold to you know all the, you know, he was a he had learjets he sold. He had he represent he had his own line of boats that he represented around the world. He had five offices around the world.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly negative, slightly dominant, slightly guarded; reads as affection, elation, longing; style: casual, conversational; average recording, quiet background; genuineness 4.7/6; vocal-burst blend 10.0/10; 23.6s, EN.
903763_00202052 · in -20.6 dBFS · gain +0.6 dB · podcast-03749
(jealousy and envy, sourness, bitterness · normal-paced, normally alert, moderately variable, casual) Well, his business had went south and he got in a sting with the DEA. He offered to sell him a plane for cash and hide the drug money, and that was illegal. So Dave got busted. So I get a phone call at my house. And oh, and in jail to make sure that I never busted any of my friends that I grew up with, my high school friends, the you know, that group of guys that are on your, you know, like your hands, like the gloves on your hand.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, slightly guarded; reads as jealousy and envy, sourness, bitterness; style: casual, monologue; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 9.2/10; 24.9s, EN.
903763_00204408 · in -22.0 dBFS · gain +2.0 dB · podcast-01532
(embarrassment, shame, disappointment · normal-paced, normally alert, fairly steady, casual) I told my best friend who was a big mouth, I go, listen, I don't know, man. Anthony, I might have to do something, work for the cops. So just tell everybody down there, you know, that (ahem) you know, don't talk to me. So I was, I took a I took a poison pill. Basically, what I made it so that when I came back from now, most guys weren't gonna talk to me anyhow, right?
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly negative, slightly dominant, slightly guarded; reads as embarrassment, shame, disappointment; style: casual, conversational; average recording, quiet background; genuineness 5.2/6; vocal-burst blend 10.0/10; 20.4s, EN.
903763_00206896 · in -22.1 dBFS · gain +2.1 dB · podcast-03738