k-PXR-k4 — voice-corrected

PXR at chain length k=4, all corpora, at the mining floor.

This is the voice-conversion-corrected version of this tier. The original, uncorrected page is still there and unchanged: t_k-PXR-k4.html. Every card below carries both renders so you can switch between them without leaving the page. What was done, and what it measured →
20chains converted
60segments re-voiced
0.783 → 0.760median worst-to-anchor identity cosine
98 %median emotion delta retained
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken.

(underlined, plain · delivery, style) — the tag before the words. Emotions first, then how it is delivered. Underlined descriptors are the ones that change across this chain — anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words are a real non-speech sound, printed where it happens.

The full generated caption for any clip is under “full caption & clip details”. Its perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.

The tags and captions describe the ORIGINAL clips — they are the corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the converted audio are the ones printed in each card's own paragraph and chips.
Emotional Numbness ↓  /  Disgustidentity −0.04 emotion 63 %   k-PXR-k4 · #1

This chain comes from the proxy rule: the same two-sided test as above, but because Disgust is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Disgust below average — 0.33, lower than 67 % of clips in this corpus — and ends with it clearly present at 0.65, higher than 65 % of clips in this corpus. That is a total rise of 0.32.

At the same time Emotional Numbness goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.60 (higher than 60 % of clips in this corpus), a change of -0.37. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.00, then +0.20, then +0.12 — a plateau around step 1, where it barely moves.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.89 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.89 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.89. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 30 s · snippets

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.808 before conversion and 0.765 after — it fell by 0.042. Neighbour-to-neighbour the worst pair went 0.808 → 0.765. (The earlier render, with segment 1 left raw, scores 0.761 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disgust moved +0.322 in the original and +0.202 after conversion — 63 % of the delta retained. On the other named axis, Emotional Numbness, -0.372 became -0.380.

Quality. Mean predicted overall quality across the segments went 2.94 → 2.97 (+0.03) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.808 → 0.765 -0.042identity cos neighbours 0.808 → 0.765d_b rescored +0.322 → +0.202d_a rescored -0.372 → -0.380d_a mined -0.368d_b mined 0.322min_cos_consec (site) 0.8910min_cos_anchor (site) 0.8877dataset snippetslang ?speaker batch88_part2_batch88_parttotal 29.4schain gain +3.2 dBseam step 1.5 dBcrossfades 100/100/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, good recording, no background noise, normal-paced, slightly relaxed, moderate pitch range
(emotional numbness, contemplation · energised, fairly steady, no disfluency, narration) at the core of Manson's beliefs was the concept of
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, neutral openness; reads as emotional numbness, contemplation; style: narration, formal; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 3.3/10; 4.0s.
batch88_part2_batch88_part2_chunk_1792_1_1464480 · in -19.8 dBFS · gain -0.2 dB · snippets-01345
(normally alert, steady, almost no disfluency, newsreading) Boko Haram, a militant Islamist group, emerged in Nigeria in the early 2000s and has since gained notoriety for its brutal tactics and a series of high-profile attacks across the region.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 12.2s.
batch88_part2_batch88_part2_chunk_1792_1_1464548 · in -21.2 dBFS · gain +1.2 dB · snippets-01345
(normally alert, steady, almost no disfluency, narration) The events also highlight the importance of recognizing and addressing the warning signs of radicalization and extremist beliefs within communities.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, newsreading; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.9/10; 8.9s.
batch88_part2_batch88_part2_chunk_1792_1_1464679 · in -21.0 dBFS · gain +1.0 dB · snippets-01345
(normally alert, fairly steady, almost no disfluency, casual) was a Japanese doomsday cult founded by Shoko Asahara in 1984.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, conversational; good recording, no background noise; genuineness 2.0/6; vocal-burst blend 2.2/10; 4.8s.
batch88_part2_batch88_part2_chunk_1792_1_1464772 · in -20.4 dBFS · gain +0.5 dB · snippets-01345
Fatigue Exhaustion ↓  /  Malevolence Maliceidentity −0.03 emotion 135 %   k-PXR-k4 · #2

This chain comes from the proxy rule: the same two-sided test as above, but because Malevolence Malice is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Malevolence Malice below average — 0.25, lower than 75 % of clips in this corpus — and ends with it strongly present at 0.84, higher than 84 % of clips in this corpus. That is a total rise of 0.59.

At the same time Fatigue Exhaustion goes the other way, from 0.79 (higher than 79 % of clips in this corpus) to 0.38 (lower than 62 % of clips in this corpus), a change of -0.41. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.19, then +0.21, then +0.19 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.84 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.81 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.84. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 25 s · zh · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.774 before conversion and 0.746 after — it fell by 0.028. Neighbour-to-neighbour the worst pair went 0.774 → 0.746. (The earlier render, with segment 1 left raw, scores 0.746 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Malevolence Malice moved +0.589 in the original and +0.792 after conversion — 135 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Fatigue Exhaustion, -0.408 became -0.142.

Quality. Mean predicted overall quality across the segments went 2.99 → 3.10 (+0.11) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.774 → 0.746 -0.028identity cos neighbours 0.774 → 0.746d_b rescored +0.589 → +0.792d_a rescored -0.408 → -0.142d_a mined -0.408d_b mined 0.589min_cos_consec (site) 0.8150min_cos_anchor (site) 0.8401dataset emolialang zhspeaker ZH_B00038_S01623total 23.5schain gain +0.7 dBseam step 1.5 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, no background noise, normally alert, slightly relaxed, fairly steady, no disfluency
(measured, monologue, didactic) 所有各方都应立即取消他们的攻击计划,并开始谈判,以解决他们彼此之间的矛盾问题。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, didactic; average recording, no background noise; genuineness 1.1/6; vocal-burst blend 3.6/10; 7.4s, ZH.
ZH_B00038_S01623_W000056 · in -17.7 dBFS · gain -2.3 dB · emolia-03655
(measured, formal, narration) 使用违禁武器的行为是不能容忍的。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, narration; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 3.8/10; 3.2s, ZH.
ZH_B00038_S01623_W000057 · in -16.5 dBFS · gain -3.5 dB · emolia-03655
(fast, formal, narration) 皇室成员没有兴趣参与调解守护者之间的类似矛盾。
full caption & clip details
An adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, narration; average recording, no background noise; genuineness 1.0/6; vocal-burst blend 3.5/10; 4.7s, ZH.
ZH_B00038_S01623_W000058 · in -17.0 dBFS · gain -3.0 dB · emolia-03655
(measured, monologue, didactic) 此外,有报告称,一些守护者准备在冲突中使用大规模杀伤性武器WMD以控制局势。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, didactic; average recording, no background noise; genuineness 0.9/6; vocal-burst blend 3.6/10; 8.8s, ZH.
ZH_B00038_S01623_W000059 · in -19.4 dBFS · gain -0.6 dB · emolia-03655
Shame ↓  /  Infatuationidentity +0.08 emotion 67 %   k-PXR-k4 · #3

This chain comes from the proxy rule: the same two-sided test as above, but because Infatuation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Infatuation clearly present — 0.73, higher than 74 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.26.

At the same time Shame goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.72 (higher than 72 % of clips in this corpus), a change of -0.28. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.23, then +0.03, then -0.00 — not a clean run: step 3 moves back the other way by 0.00 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.66 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.65 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.66, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 38 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.676 before conversion and 0.753 after — it rose by 0.077. Neighbour-to-neighbour the worst pair went 0.604 → 0.635. (The earlier render, with segment 1 left raw, scores 0.694 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Infatuation moved +0.260 in the original and +0.174 after conversion — 67 % of the delta retained. On the other named axis, Shame, -0.278 became -0.159.

Quality. Mean predicted overall quality across the segments went 2.92 → 3.05 (+0.13) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.676 → 0.753 +0.077identity cos neighbours 0.604 → 0.635d_b rescored +0.260 → +0.174d_a rescored -0.278 → -0.159d_a mined -0.278d_b mined 0.260min_cos_consec (site) 0.6452min_cos_anchor (site) 0.6646dataset emolialang enspeaker EN_B00030_S08249total 37.2schain gain +3.1 dBseam step 1.0 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an elderly masculine voice · no background noise, very low-energy, steady
(shame, sadness, bitterness · measured, slightly relaxed, little disfluency, whispered) I wedded, nor dreaded the curse I had invoked, and its bitterness was not visited upon me at once.
full caption & clip details
An elderly masculine voice; delivery is very low-energy, measured, slightly relaxed, steady; timbre is slightly warm, slightly dark, rough, very full; clear, little disfluency, fairly narrow pitch, light breath; affect is mildly negative, neutral stance, fairly guarded; reads as shame, sadness, bitterness; style: whispered, narration; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 3.6/10; 7.3s, EN.
EN_B00030_S08249_W000333 · in -23.9 dBFS · gain +3.9 dB · emolia-00838
(contentment, awe, contemplation · slow, relaxed, little disfluency, whispered) But once again, in the silence of the night, they came through my lattice the soft sighs which had forsaken me, and they modelled themselves into a familiar sweet voice, saying
full caption & clip details
An elderly masculine voice; delivery is very low-energy, slow, relaxed, steady; timbre is warm, slightly dark, slightly rough, balanced body; slurred, little disfluency, fairly narrow pitch, minimal breath; affect is mildly negative, neutral stance, neutral openness; reads as contentment, awe, contemplation; style: whispered, ASMR; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 1.9/10; 12.4s, EN.
EN_B00030_S08249_W000334 · in -25.3 dBFS · gain +5.3 dB · emolia-00838
(contentment, longing, infatuation · slow, relaxed, no disfluency, whispered) Sleep in peace. For the spirit of love.
full caption & clip details
An elderly masculine voice; delivery is very low-energy, slow, relaxed, steady; timbre is slightly warm, dark, smooth, very full; somewhat unclear, no disfluency, fairly narrow pitch, light breath; affect is mildly positive, submissive, neutral openness; reads as contentment, longing, infatuation; style: whispered, monologue; good recording, no background noise; genuineness 1.4/6; vocal-burst blend 7.6/10; 3.3s, EN.
EN_B00030_S08249_W000335 · in -24.2 dBFS · gain +4.2 dB · emolia-00838
(infatuation, sexual lust, awe · slow, relaxed, little disfluency, whispered) In taking to thy passionate heart her who is um, who is Ermengarde, thou art absolved for reasons which shall be made known to thee in heaven of thy vows unto Elenora.
full caption & clip details
An elderly masculine voice; delivery is very low-energy, slow, relaxed, steady; timbre is slightly warm, slightly dark, slightly rough, very full; slurred, little disfluency, narrow pitch range, minimal breath; affect is mildly negative, submissive, neutral openness; reads as infatuation, sexual lust, awe; style: whispered, ASMR; average recording, no background noise; genuineness 0.5/6; vocal-burst blend 0.6/10; 14.9s, EN.
EN_B00030_S08249_W000336 · in -24.6 dBFS · gain +4.6 dB · emolia-00838
Relief ↓  /  Jealousy and Envyidentity +0.50 emotion 73 %   k-PXR-k4 · #4

This chain comes from the proxy rule: the same two-sided test as above, but because Jealousy and Envy is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Jealousy and Envy clearly present — 0.66, higher than 66 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.33.

At the same time Relief goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.72 (higher than 72 % of clips in this corpus), a change of -0.27. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.12, then +0.13, then +0.08 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.14 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.13 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.14, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 47 s · en · podcast

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.059 before conversion and 0.562 after — it rose by 0.504. Neighbour-to-neighbour the worst pair went 0.059 → 0.562. (The earlier render, with segment 1 left raw, scores 0.489 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Jealousy and Envy moved +0.336 in the original and +0.244 after conversion — 73 % of the delta retained, which is most of it. On the other named axis, Relief, -0.310 became -0.515.

Quality. Mean predicted overall quality across the segments went 2.89 → 3.10 (+0.21) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.059 → 0.562 +0.504identity cos neighbours 0.059 → 0.562d_b rescored +0.336 → +0.244d_a rescored -0.310 → -0.515d_a mined -0.275d_b mined 0.330min_cos_consec (site) 0.1321min_cos_anchor (site) 0.1410dataset podcastlang enspeaker 946244total 46.4schain gain +2.6 dBseam step 2.2 dBcrossfades 150/150/100 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · neutral-toned, fairly smooth, normal-paced, some disfluency, average clarity
(relief, triumph, pride · normally alert, neutral tension, moderately variable, casual) I enjoyed watching it. I was extremely happy when I was like, Oh, wow. (ahem) Uh finally, I don't have to see the Rangers all over my timeline in the postseason anymore. Like it's done. I've been released. Yep.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, vulnerable; reads as relief, triumph, pride; style: casual, dramatic; good recording, no background noise; genuineness 3.8/6; vocal-burst blend 4.1/10; 14.1s, EN.
946244_00010608 · in -23.4 dBFS · gain +3.4 dB · podcast-04301
(teasing, fear, relief · normally alert, slightly relaxed, fairly steady, casual) And I don't think you're gonna have to see the Rangers go that far this year either.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as teasing, fear, relief; style: casual, conversational; average recording, no background noise; genuineness 3.4/6; vocal-burst blend 3.1/10; 3.6s, EN.
946244_00012024 · in -22.0 dBFS · gain +2.0 dB · podcast-04297
(disappointment, bitterness, interest · very low-energy, slightly relaxed, fairly steady, casual) It's because they're the Rangers. The closest they've gotten in the last few years was that loss to the Kings, in which Alex m Alec Martinez took care of them in that three two victory, and I don't think they'll be able to overcome that for A few more years they had to let Terasenko walk. Not sure if Kane's gonna play. Kane might be retired with all his injuries.
full caption & clip details
A young adult masculine voice; delivery is very low-energy, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as disappointment, bitterness, interest; style: casual, monologue; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 5.8/10; 19.8s, EN.
946244_00012552 · in -22.3 dBFS · gain +2.3 dB · podcast-06371
(jealousy and envy, emotional numbness, doubt · normally alert, slightly relaxed, fairly steady, conversational) Has there been any update on a Patrick Kane contract? Anything surrounding him? I've seen a lot of speculation on Twitter about him, but nothing's ever been confirmed.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as jealousy and envy, emotional numbness, doubt; style: conversational, casual; good recording, no background noise; mildly explicit content; genuineness 3.4/6; vocal-burst blend 3.7/10; 9.4s, EN.
946244_00014551 · in -22.5 dBFS · gain +2.5 dB · podcast-04293
Emotional Numbness ↓  /  Sexual Lustidentity +0.04 emotion 106 %   k-PXR-k4 · #5

This chain comes from the proxy rule: the same two-sided test as above, but because Sexual Lust is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Sexual Lust clearly present — 0.58, higher than 58 % of clips in this corpus — and ends with it strongly present at 0.90, higher than 90 % of clips in this corpus. That is a total rise of 0.31.

At the same time Emotional Numbness goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.71 (higher than 71 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.11, then +0.23, then -0.03 — not a clean run: step 3 moves back the other way by 0.03 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.81 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.82 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.81. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 42 s · en · podcast

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.769 before conversion and 0.810 after — it rose by 0.041. Neighbour-to-neighbour the worst pair went 0.833 → 0.778. (The earlier render, with segment 1 left raw, scores 0.726 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sexual Lust moved +0.311 in the original and +0.328 after conversion — 106 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Emotional Numbness, -0.266 became -0.238.

Quality. Mean predicted overall quality across the segments went 3.12 → 3.27 (+0.15) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.769 → 0.810 +0.041identity cos neighbours 0.833 → 0.778d_b rescored +0.311 → +0.328d_a rescored -0.266 → -0.238d_a mined -0.262d_b mined 0.313min_cos_consec (site) 0.8152min_cos_anchor (site) 0.8099dataset podcastlang enspeaker 635231total 41.4schain gain +3.0 dBseam step 3.6 dBcrossfades 150/150/150 ms
Script — 4 chunks, 3 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-bright
(emotional numbness, helplessness, disappointment · measured, very low-energy, slightly relaxed, monologue) not positive. And we're just slipping further and further behind the English Premier League
full caption & clip details
An adult masculine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, normal breath; affect is mildly negative, neutral stance, neutral openness; reads as emotional numbness, helplessness, disappointment; style: monologue, whispered; good recording, no background noise; genuineness 2.1/6; vocal-burst blend 3.5/10; 6.8s, EN.
635231_00049160 · in -32.4 dBFS · gain +12.4 dB · podcast-04191
(fatigue exhaustion, emotional numbness, infatuation · measured, subdued, relaxed, monologue) (low mumble) and other leagues in Europe. And, you know, to stay on the kind of Celtic and Rangers comparison again, you know, back in the day, you know,
full caption & clip details
An adult masculine voice; delivery is subdued, measured, relaxed, fairly steady; timbre is slightly warm, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, normal breath; affect is mildly negative, neutral stance, neutral openness; reads as fatigue exhaustion, emotional numbness, infatuation; style: monologue, whispered; good recording, no background noise; genuineness 1.9/6; vocal-burst blend 4.4/10; 9.2s, EN.
635231_00049840 · in -31.7 dBFS · gain +11.7 dB · podcast-04149
(jealousy and envy, sexual lust, infatuation · measured, normally alert, slightly relaxed, narration) Celtic and Rangers used to compare themselves to kind of top kind of (low mumble) um European clubs, you know, your Ajax, Porto, that used to be kind of the level. That was the peer group. And now it's more like your Copenhagen's.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as jealousy and envy, sexual lust, infatuation; style: narration, whispered; good recording, no background noise; mildly explicit content; genuineness 2.1/6; vocal-burst blend 4.8/10; 12.3s, EN.
635231_00050760 · in -31.3 dBFS · gain +11.3 dB · podcast-04137
(slow, very low-energy, neutral tension, casual) (low mumble) Um, your Club Bruges. And but actually, you know, the m longer this goes on, the the lower that quality of club is. And it's such a slow (contented sigh) sometimes these things happen so slowly.
full caption & clip details
An adult masculine voice; delivery is very low-energy, slow, neutral tension, moderately variable; timbre is slightly warm, neutral-bright, slightly rough, full; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly negative, slightly dominant, neutral openness; no dominant emotion; style: casual, monologue; average recording, quiet background; mildly explicit content; genuineness 3.6/6; vocal-burst blend 4.2/10; 13.5s, EN.
635231_00052040 · in -29.1 dBFS · gain +9.1 dB · podcast-03756
Contemplation ↓  /  Affectionidentity −0.05 emotion 140 %   k-PXR-k4 · #6

This chain comes from the proxy rule: the same two-sided test as above, but because Affection is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Affection clearly present — 0.60, higher than 60 % of clips in this corpus — and ends with it strongly present at 0.85, higher than 85 % of clips in this corpus. That is a total rise of 0.25.

At the same time Contemplation goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.60 (higher than 60 % of clips in this corpus), a change of -0.39. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.00, then +0.24, then +0.02 — a plateau around step 1, where it barely moves.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.82 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.82 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.82. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 29 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.785 before conversion and 0.737 after — it fell by 0.048. Neighbour-to-neighbour the worst pair went 0.782 → 0.708. (The earlier render, with segment 1 left raw, scores 0.548 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Affection moved +0.252 in the original and +0.353 after conversion — 140 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Contemplation, -0.391 became -0.386.

Quality. Mean predicted overall quality across the segments went 2.52 → 2.86 (+0.34) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.785 → 0.737 -0.048identity cos neighbours 0.782 → 0.708d_b rescored +0.252 → +0.353d_a rescored -0.391 → -0.386d_a mined -0.391d_b mined 0.252min_cos_consec (site) 0.8185min_cos_anchor (site) 0.8181dataset emolialang enspeaker EN_wYThp1HpUwutotal 27.7schain gain +2.5 dBseam step 1.6 dBcrossfades 150/150/100 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · fairly smooth, slightly relaxed
(contemplation, concentration, interest · normal-paced, normally alert, fairly steady, monologue) The next question I'd like you to think about is what is the problem of the story and how did that problem get solved?
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as contemplation, concentration, interest; style: monologue, ASMR; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 0.1/10; 6.8s, EN.
EN_wYThp1HpUwu_W000163 · in -21.4 dBFS · gain +1.4 dB · emolia-00925
(measured, normally alert, fairly steady, monologue) When I've asked this question of other students who have heard the story,
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is slightly warm, neutral-bright, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; no dominant emotion; style: monologue, casual; average recording, quiet background; genuineness 2.7/6; vocal-burst blend 1.0/10; 4.5s, EN.
EN_wYThp1HpUwu_W000164 · in -17.0 dBFS · gain -3.0 dB · emolia-00925
(sadness, distress, disappointment · normal-paced, normally alert, steady, formal) Other kids often say that the problem of the story was that it was really difficult to get to the North Pole.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as sadness, distress, disappointment; style: formal, monologue; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 0.5/10; 6.3s, EN.
EN_wYThp1HpUwu_W000165 · in -17.0 dBFS · gain -3.0 dB · emolia-00925
(slow, very low-energy, fairly steady, ASMR) And the way that it was solved was with (low mumble) Matthew Henson's, uh, (low mumble) his, (low mumble) uh, resourcefulness and the rest of the group's teamwork.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, slow, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, slightly thin; average clarity, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly positive, submissive, neutral openness; no dominant emotion; style: ASMR, monologue; average recording, quiet background; genuineness 2.6/6; vocal-burst blend 0.0/10; 10.6s, EN.
EN_wYThp1HpUwu_W000166 · in -15.6 dBFS · gain -4.4 dB · emolia-00925
Emotional Numbness ↓  /  Disgustidentity −0.08 emotion 103 %   k-PXR-k4 · #7

This chain comes from the proxy rule: the same two-sided test as above, but because Disgust is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Disgust below average — 0.33, lower than 67 % of clips in this corpus — and ends with it strongly present at 0.76, higher than 76 % of clips in this corpus. That is a total rise of 0.44.

At the same time Emotional Numbness goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.69 (higher than 69 % of clips in this corpus), a change of -0.30. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.00, then +0.20, then +0.23 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.94 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.96 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.94), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 38 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.927 before conversion and 0.847 after — it fell by 0.081. Neighbour-to-neighbour the worst pair went 0.833 → 0.771. (The earlier render, with segment 1 left raw, scores 0.509 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disgust moved +0.436 in the original and +0.449 after conversion — 103 % of the delta retained, which is essentially all of it. On the other named axis, Emotional Numbness, -0.298 became -0.287.

Quality. Mean predicted overall quality across the segments went 2.96 → 3.07 (+0.11) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.927 → 0.847 -0.081identity cos neighbours 0.833 → 0.771d_b rescored +0.436 → +0.449d_a rescored -0.298 → -0.287d_a mined -0.298d_b mined 0.436min_cos_consec (site) 0.9601min_cos_anchor (site) 0.9387dataset emolialang enspeaker EN_Ps4Ps7rMTwktotal 36.7schain gain +1.8 dBseam step 1.4 dBcrossfades 100/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(emotional numbness · steady, formal, newsreading) Mechanical engineering, welding, electrical machinery, control systems, electric circuits, engine room simulators and graphics
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: formal, newsreading; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.1/10; 7.4s, EN.
EN_Ps4Ps7rMTwk_W000094 · in -15.8 dBFS · gain -4.2 dB · emolia-00872
(fairly steady, newsreading, formal) Bronze statues of Fulton and Christopher Columbus represent commerce on the balustrade of the galleries of the main Reading Room in the Thomas Jefferson Building of the Library of Congress on Capitol Hill in Washington, D.C.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: newsreading, formal; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.0/10; 11.6s, EN.
EN_Ps4Ps7rMTwk_W000095 · in -15.3 dBFS · gain -4.7 dB · emolia-00872
(awe · fairly steady, formal, authoritative) They are two of sixteen historical figures, each pair representing one of the eight pillars of civilization
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as awe; style: formal, authoritative; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.0/10; 5.7s, EN.
EN_Ps4Ps7rMTwk_W000096 · in -14.1 dBFS · gain -5.9 dB · emolia-00872
(fairly steady, newsreading, formal) The Guatemalan government in 1910 erected a bust of Fulton in one of the parks of Guatemala City.In 2006, he was inducted into the
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 12.7s, EN.
EN_Ps4Ps7rMTwk_W000097 · in -15.0 dBFS · gain -5.0 dB · emolia-00872
Triumph ↓  /  Intoxication Altered States of Consciousnessidentity −0.07 emotion 74 %   k-PXR-k4 · #8

This chain comes from the proxy rule: the same two-sided test as above, but because Intoxication Altered States of Consciousness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Intoxication Altered States of Consciousness around average — 0.47, lower than 53 % of clips in this corpus — and ends with it strongly present at 0.82, higher than 82 % of clips in this corpus. That is a total rise of 0.35.

At the same time Triumph goes the other way, from 0.67 (higher than 67 % of clips in this corpus) to 0.38 (lower than 62 % of clips in this corpus), a change of -0.28. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.16, then +0.15, then +0.03 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.93 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.93 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.93), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 27 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.916 before conversion and 0.849 after — it fell by 0.067. Neighbour-to-neighbour the worst pair went 0.913 → 0.840. (The earlier render, with segment 1 left raw, scores 0.739 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Intoxication Altered States of Consciousness moved +0.350 in the original and +0.257 after conversion — 74 % of the delta retained, which is most of it. On the other named axis, Triumph, -0.283 became -0.192.

Quality. Mean predicted overall quality across the segments went 2.99 → 3.04 (+0.05) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.916 → 0.849 -0.067identity cos neighbours 0.913 → 0.840d_b rescored +0.350 → +0.257d_a rescored -0.283 → -0.192d_a mined -0.283d_b mined 0.350min_cos_consec (site) 0.9312min_cos_anchor (site) 0.9254dataset emolialang enspeaker EN_RxWaYFXy1bItotal 26.0schain gain +0.8 dBseam step 1.7 dBcrossfades 100/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · neutral-toned, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert, slightly relaxed
(fairly steady, formal, casual) August 16, 1945 The Nakajima Aircraft Company changed its name to Fuji-Sonyo Co., Ltd.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, casual; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.1/10; 8.1s, EN.
EN_RxWaYFXy1bI_W000025 · in -14.7 dBFS · gain -5.3 dB · emolia-00596
(steady, formal, monologue) November 6, 1945 – The GHQ defined Fuji-Sonyo as the Zebatsu and decided to disband them
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 8.2s, EN.
EN_RxWaYFXy1bI_W000026 · in -16.0 dBFS · gain -4.0 dB · emolia-00596
(steady, monologue, formal) May 1950 Fuji-San-Yoko, Limited was disbanded
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 2.6/10; 4.4s, EN.
EN_RxWaYFXy1bI_W000027 · in -13.6 dBFS · gain -6.4 dB · emolia-00596
(fairly steady, formal, monologue) July 1950 Fuji-Sung Yoko, Ltd. was divided into 12 companies
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 1.2/10; 5.9s, EN.
EN_RxWaYFXy1bI_W000028 · in -14.2 dBFS · gain -5.8 dB · emolia-00596
Concentration ↓  /  Interestidentity +0.57 emotion 104 %   k-PXR-k4 · #9

This chain comes from the proxy rule: the same two-sided test as above, but because Interest is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Interest clearly present — 0.67, higher than 67 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.30.

At the same time Concentration goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.69 (higher than 69 % of clips in this corpus), a change of -0.30. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.01, then +0.11, then +0.18 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.33 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.31 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.33, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 60 s · da · eurospeech

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.329 before conversion and 0.899 after — it rose by 0.570. Neighbour-to-neighbour the worst pair went 0.292 → 0.880. (The earlier render, with segment 1 left raw, scores 0.811 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Interest moved +0.300 in the original and +0.312 after conversion — 104 % of the delta retained, which is essentially all of it. On the other named axis, Concentration, -0.296 became -0.302.

Quality. Mean predicted overall quality across the segments went 3.10 → 3.40 (+0.30) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.329 → 0.899 +0.570identity cos neighbours 0.292 → 0.880d_b rescored +0.300 → +0.312d_a rescored -0.296 → -0.302d_a mined -0.296d_b mined 0.299min_cos_consec (site) 0.3077min_cos_anchor (site) 0.3314dataset eurospeechlang daspeaker denmark_20181M077_2019-03-total 58.6schain gain +1.9 dBseam step 1.0 dBcrossfades 150/150/150 ms
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: a middle-aged masculine voice · neutral-toned, balanced body, average recording, measured, slightly relaxed
(concentration, disappointment, jealousy and envy · subdued, steady, little disfluency, monologue) Jeg synes, at grisene burde have deres egen lov. Efter det her lovforslag har hundene fortsat deres egen lov – jeg synes faktisk, vi skulle fastholde, at de 30 millioner grise også skulle have en lov. På den baggrund kan Enhedslisten ikke støtte lovforslaget, men vi vil meget gerne være med til at arbejde for mere dyrevelfærd.
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, disappointment, jealousy and envy; style: monologue, formal; average recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.3/10; 17.1s, DA.
denmark_20181M077_2019-03-28_1000_15950048_15967136 · in -25.1 dBFS · gain +5.1 dB · eurospeech-00360
(concentration, thankfulness gratitude · normally alert, fairly steady, frequent disfluency, monologue) Tak for det. Jeg har bare et enkelt spørgsmål. Ordføreren nævnte, at vi har en alt for stor animalsk produktion i Danmark. Hvor meget skal den reduceres efter Enhedslistens opfattelse? Er det en fjernelse af den eksportorienterede del,
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, thankfulness gratitude; style: monologue, didactic; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 0.2/10; 16.2s, DA.
denmark_20181M077_2019-03-28_1000_15967136_15983360 · in -25.9 dBFS · gain +5.8 dB · eurospeech-00360
(triumph, pride, concentration · normally alert, fairly steady, some disfluency, monologue) (ahem) I øjeblikket bruger man 80 pct. af landbrugsarealet til den animalske produktion. Vi går ind for, at man reducerer landbrugsarealet med 500.000 ha, så vi får 200.000 ha mere natur, 100.000 ha mere skov,
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as triumph, pride, concentration; style: monologue, didactic; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 2.8/10; 13.1s, DA.
denmark_20181M077_2019-03-28_1000_15995920_16008976 · in -24.2 dBFS · gain +4.2 dB · eurospeech-00360
(interest · normally alert, fairly steady, some disfluency, monologue) og gerne omstiller (low mumble) til nogle energiafgrøder (low mumble) i landbruget, og det kunne så fortsat være landbrugsarealer. Vi ser gerne, at der er nogle organogene jorder, 100.000 ha, som tages ud af omdrift, og det kan så stadig væk
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as interest; style: monologue; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 5.8/10; 12.8s, DA.
denmark_20181M077_2019-03-28_1000_16008976_16021776 · in -26.2 dBFS · gain +6.2 dB · eurospeech-00360
Concentration ↓  /  Infatuationidentity −0.12 emotion 83 %   k-PXR-k4 · #10

This chain comes from the proxy rule: the same two-sided test as above, but because Infatuation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Infatuation around average — 0.51, right about the corpus median — and ends with it strongly present at 0.88, higher than 88 % of clips in this corpus. That is a total rise of 0.37.

At the same time Concentration goes the other way, from 0.93 (higher than 93 % of clips in this corpus) to 0.67 (higher than 67 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.14, then +0.10, then +0.12 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.96 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.95 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.96), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 45 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.880 before conversion and 0.755 after — it fell by 0.125. Neighbour-to-neighbour the worst pair went 0.913 → 0.826. (The earlier render, with segment 1 left raw, scores 0.584 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Infatuation moved +0.364 in the original and +0.303 after conversion — 83 % of the delta retained, which is most of it. On the other named axis, Concentration, -0.264 became -0.299.

Quality. Mean predicted overall quality across the segments went 3.07 → 3.18 (+0.12) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.880 → 0.755 -0.125identity cos neighbours 0.913 → 0.826d_b rescored +0.364 → +0.303d_a rescored -0.264 → -0.299d_a mined -0.264d_b mined 0.369min_cos_consec (site) 0.9535min_cos_anchor (site) 0.9551dataset emolialang enspeaker EN_kp2oAu11ts8total 44.0schain gain +0.6 dBseam step 1.1 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(concentration · steady, almost no disfluency, newsreading, formal) As the same compiler is available for all of the above operating systems, there is no need for recoding to produce identical products for different platforms, except when operating system dependent features are used. Cross compiling is supported with Ming-W.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 15.8s, EN.
EN_kp2oAu11ts8_W000022 · in -14.8 dBFS · gain -5.2 dB · emolia-02415
(fairly steady, almost no disfluency, authoritative, newsreading) Under Microsoft Windows, Harbor is more stable but less well documented than Clipper, but has multi-platform capability and is more transparent, customizable and can run from a USB flash drive
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: authoritative, newsreading; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.0/10; 12.4s, EN.
EN_kp2oAu11ts8_W000023 · in -15.0 dBFS · gain -5.0 dB · emolia-02415
(fairly steady, no disfluency, formal, monologue) Under Linux and Windows Mobile, Clipper source code can be compiled with Harbor with very little adaptation
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 6.5s, EN.
EN_kp2oAu11ts8_W000024 · in -13.9 dBFS · gain -6.1 dB · emolia-02415
(steady, no disfluency, formal, newsreading) Most software originally written to run on XBase++, Flagship, FoxPro, X Harbor and others dialects can be compiled with Harbor with some adaptation
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 9.8s, EN.
EN_kp2oAu11ts8_W000025 · in -15.1 dBFS · gain -4.9 dB · emolia-02415
Thankfulness Gratitude ↓  /  Elationidentity +0.31 emotion 309 %   k-PXR-k4 · #11

This chain comes from the proxy rule: the same two-sided test as above, but because Elation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Elation clearly present — 0.69, higher than 69 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.26.

At the same time Thankfulness Gratitude goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.70 (higher than 70 % of clips in this corpus), a change of -0.27. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.18, then +0.11, then -0.03 — not a clean run: step 3 moves back the other way by 0.03 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.43 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.33 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.43, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 48 s · es · podcast

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.495 before conversion and 0.809 after — it rose by 0.314. Neighbour-to-neighbour the worst pair went 0.361 → 0.809. (The earlier render, with segment 1 left raw, scores 0.702 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Elation moved +0.264 in the original and +0.818 after conversion — 309 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Thankfulness Gratitude, -0.259 became -0.654.

Quality. Mean predicted overall quality across the segments went 2.52 → 3.15 (+0.63) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.495 → 0.809 +0.314identity cos neighbours 0.361 → 0.809d_b rescored +0.264 → +0.818d_a rescored -0.259 → -0.654d_a mined -0.271d_b mined 0.261min_cos_consec (site) 0.3341min_cos_anchor (site) 0.4271dataset podcastlang esspeaker 115744total 46.6schain gain +2.0 dBseam step 2.3 dBcrossfades 150/150/100 ms
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · fast, neutral tension, moderately variable, wide pitch range
(thankfulness gratitude, embarrassment, amusement · normally alert, frequent disfluency, somewhat unclear, casual) su opinión. Y si encima todo esto le juntas con gente que no juega en serio, que se lo está pasando, vamos, que empiezan con coña y tal, pues fue muy divertido.
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, thin; somewhat unclear, frequent disfluency, wide pitch range, audible breath; affect is positive, slightly dominant, slightly guarded; reads as thankfulness gratitude, embarrassment, amusement; style: casual, playful; average recording, quiet background; genuineness 5.0/6; vocal-burst blend 8.8/10; 10.8s, ES.
115744_00675560 · in -22.6 dBFS · gain +2.6 dB · podcast-05381
(relief, disappointment, pride · normally alert, frequent disfluency, somewhat unclear, casual) Pues la vez. Y no sé, impresiones de esto, pues que estuvo divertido. La verdad es que yo no sobreviví. Me llevé a uno por delante. (low mumble) Y no sobreviví.
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, neutral tension, moderately variable; timbre is neutral-toned, very dark, slightly rough, thin; somewhat unclear, frequent disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, slightly guarded; reads as relief, disappointment, pride; style: casual, playful; below-average recording, some background noise; genuineness 5.2/6; vocal-burst blend 9.3/10; 10.8s, ES.
115744_00676736 · in -21.3 dBFS · gain +1.3 dB · podcast-05387
(elation, embarrassment, pride · energised, some disfluency, average clarity, casual) Sí, no, nosotros hicimos una conga y quitándome a mí que yo me aparte un poco del fregado y me metí al asteroide para pillar unos preciosos torpedos de plasma. Torpedos de plasma que...
full caption & clip details
A young adult masculine voice; delivery is energised, fast, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as elation, embarrassment, pride; style: casual, dramatic; average recording, quiet background; genuineness 4.7/6; vocal-burst blend 8.2/10; 9.2s, ES.
115744_00690656 · in -22.9 dBFS · gain +2.9 dB · podcast-05397
(elation · energised, some disfluency, somewhat unclear, casual) Bueno, pues con esta. Sí, con esta primera fase, pues al final, pues por puntos, se hicieron los cruces para la seconda fase. A 100 (surprised gasp) puntos, pero la lista tenía que ser, como bien hemos dicho antes, una lista de solo tres naves.
full caption & clip details
A young adult masculine voice; delivery is energised, fast, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as elation; style: casual, dramatic; below-average recording, some background noise; genuineness 4.8/6; vocal-burst blend 10.0/10; 16.2s, ES.
115744_00692552 · in -22.5 dBFS · gain +2.5 dB · podcast-00668
Emotional Numbness ↓  /  Thankfulness Gratitudeidentity −0.01 emotion 179 %   k-PXR-k4 · #12

This chain comes from the proxy rule: the same two-sided test as above, but because Thankfulness Gratitude is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Thankfulness Gratitude clearly present — 0.62, higher than 62 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.36.

At the same time Emotional Numbness goes the other way, from 0.87 (higher than 87 % of clips in this corpus) to 0.49 (right about the corpus median), a change of -0.38. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.23, then -0.07, then +0.20 — not a clean run: step 2 moves back the other way by 0.07 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.85 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.88 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.85. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 55 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.843 before conversion and 0.831 after — it fell by 0.011. Neighbour-to-neighbour the worst pair went 0.842 → 0.787. (The earlier render, with segment 1 left raw, scores 0.659 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Thankfulness Gratitude moved +0.363 in the original and +0.652 after conversion — 179 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Emotional Numbness, -0.380 became -0.461.

Quality. Mean predicted overall quality across the segments went 3.12 → 3.22 (+0.10) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.843 → 0.831 -0.011identity cos neighbours 0.842 → 0.787d_b rescored +0.363 → +0.652d_a rescored -0.380 → -0.461d_a mined -0.380d_b mined 0.363min_cos_consec (site) 0.8804min_cos_anchor (site) 0.8548dataset emolialang enspeaker EN_B00040_S03013total 54.0schain gain +2.6 dBseam step 1.5 dBcrossfades 100/100/100 ms
Script — 4 chunks, 3 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · fairly smooth, balanced body, energised, average clarity, wide pitch range
(normal-paced, slightly relaxed, fairly steady, casual) Caught a slant was able to bring it in for a touchdown, but also the backup quarterback PJ Walker really highlights as a low light. Had a couple interceptions throughout this training camp.
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; no dominant emotion; style: casual, storytelling; good recording, no background noise; genuineness 2.6/6; vocal-burst blend 2.5/10; 11.6s, EN.
EN_B00040_S03013_W000032 · in -24.0 dBFS · gain +4.0 dB · emolia-01016
(disappointment, shame, fatigue exhaustion · normal-paced, neutral tension, moderately variable, casual) Four passes and poor decisions today and really didn't look to be a backup at like the quality backup that we thought we were getting. Yet again, one day, one practice, just calling it out that he didn't look too great. (low mumble) Um, also I guess a highlight.
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, frequent disfluency, wide pitch range, normal breath; affect is mildly positive, slightly dominant, neutral openness; reads as disappointment, shame, fatigue exhaustion; style: casual, conversational; average recording, quiet background; genuineness 4.7/6; vocal-burst blend 6.5/10; 16.9s, EN.
EN_B00040_S03013_W000033 · in -24.1 dBFS · gain +4.1 dB · emolia-01016
(disgust, amusement, astonishment surprise · brisk, slightly tense, moderately variable, casual) (ahem) Uh, cheese claypool pancake TJ Edwards. And when he did that, he went over him and said, go to sleep and flex time. That's, that's pretty awesome.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, slightly tense, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, fairly guarded; reads as disgust, amusement, astonishment surprise; style: casual, dramatic; average recording, quiet background; mildly explicit content; genuineness 2.6/6; vocal-burst blend 1.4/10; 8.5s, EN.
EN_B00040_S03013_W000034 · in -22.3 dBFS · gain +2.3 dB · emolia-01016
(thankfulness gratitude, embarrassment, amusement · brisk, neutral tension, moderately variable, casual) (ahem) At the same time, dude, it's practice. Like I appreciate it, but let's not injure our linebacker out here. So I appreciate the frustration. And he was probably frustrated because the offense wasn't doing anything, but still like the tenacity. And that's what he brings as a former tight end that does play wide receiver.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as thankfulness gratitude, embarrassment, amusement; style: casual, playful; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 4.1/10; 17.6s, EN.
EN_B00040_S03013_W000035 · in -22.3 dBFS · gain +2.3 dB · emolia-01016
Fatigue Exhaustion ↓  /  Thankfulness Gratitudeidentity −0.03 emotion 112 %   k-PXR-k4 · #13

This chain comes from the proxy rule: the same two-sided test as above, but because Thankfulness Gratitude is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Thankfulness Gratitude around average — 0.46, lower than 54 % of clips in this corpus — and ends with it strongly present at 0.88, higher than 88 % of clips in this corpus. That is a total rise of 0.43.

At the same time Fatigue Exhaustion goes the other way, from 0.89 (higher than 89 % of clips in this corpus) to 0.42 (lower than 58 % of clips in this corpus), a change of -0.46. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.03, then +0.23, then +0.17 — a plateau around step 1, where it barely moves.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.47 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.47 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.47, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 22 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.522 before conversion and 0.494 after — it fell by 0.028. Neighbour-to-neighbour the worst pair went 0.522 → 0.494. (The earlier render, with segment 1 left raw, scores 0.517 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Thankfulness Gratitude moved +0.458 in the original and +0.512 after conversion — 112 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Fatigue Exhaustion, -0.464 became -0.433.

Quality. Mean predicted overall quality across the segments went 2.36 → 2.73 (+0.37) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.522 → 0.494 -0.028identity cos neighbours 0.522 → 0.494d_b rescored +0.458 → +0.512d_a rescored -0.464 → -0.433d_a mined -0.464d_b mined 0.427min_cos_consec (site) 0.4674min_cos_anchor (site) 0.4674dataset emolialang enspeaker EN_PpJGxcVW_IEtotal 20.7schain gain +5.3 dBseam step 3.5 dBcrossfades 100/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, fairly smooth, normally alert, slightly relaxed, moderate pitch range, light breath
(normal-paced, fairly steady, some disfluency, casual) If you've got humans in your image, got one here, quite far away, you can actually...
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, dark, fairly smooth, thin; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; no dominant emotion; style: casual, storytelling; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 3.7/10; 3.8s, EN.
EN_PpJGxcVW_IE_W000069 · in -16.3 dBFS · gain -3.7 dB · emolia-01183
(awe, pain · measured, steady, frequent disfluency, casual) This tool is brilliant. You can actually...
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as awe, pain; style: casual, monologue; good recording, no background noise; genuineness 2.3/6; vocal-burst blend 1.7/10; 3.0s, EN.
EN_PpJGxcVW_IE_W000071 · in -17.3 dBFS · gain -2.7 dB · emolia-01183
(relief · normal-paced, fairly steady, some disfluency, monologue) Take the sky reflection and put it in the water, because it would look quite strange if you didn't have the sky reflection in nice smooth water.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as relief; style: monologue, didactic; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 1.7/10; 7.4s, EN.
EN_PpJGxcVW_IE_W000072 · in -18.7 dBFS · gain -1.3 dB · emolia-01183
(normal-paced, fairly steady, little disfluency, monologue) And for landscape photographers again, you'll love this. You can actually add a water blur in. Once again, we have no water, but you can bring this up.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: monologue, casual; good recording, no background noise; genuineness 1.8/6; vocal-burst blend 3.0/10; 7.1s, EN.
EN_PpJGxcVW_IE_W000073 · in -17.3 dBFS · gain -2.7 dB · emolia-01183
Concentration ↓  /  Confusionidentity −0.04 emotion 92 %   k-PXR-k4 · #14

This chain comes from the proxy rule: the same two-sided test as above, but because Confusion is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Confusion below average — 0.34, lower than 66 % of clips in this corpus — and ends with it clearly present at 0.74, higher than 74 % of clips in this corpus. That is a total rise of 0.40.

At the same time Concentration goes the other way, from 0.88 (higher than 88 % of clips in this corpus) to 0.59 (higher than 59 % of clips in this corpus), a change of -0.29. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.15, then +0.00, then +0.25 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.90 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.92 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.90), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 35 s · zh · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.877 before conversion and 0.841 after — it fell by 0.036. Neighbour-to-neighbour the worst pair went 0.819 → 0.789. (The earlier render, with segment 1 left raw, scores 0.769 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Confusion moved +0.397 in the original and +0.364 after conversion — 92 % of the delta retained, which is essentially all of it. On the other named axis, Concentration, -0.288 became -0.296.

Quality. Mean predicted overall quality across the segments went 3.04 → 3.18 (+0.14) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.877 → 0.841 -0.036identity cos neighbours 0.819 → 0.789d_b rescored +0.397 → +0.364d_a rescored -0.288 → -0.296d_a mined -0.288d_b mined 0.397min_cos_consec (site) 0.9182min_cos_anchor (site) 0.9020dataset emolialang zhspeaker ZH_B00009_S06684total 34.3schain gain +0.6 dBseam step 1.1 dBcrossfades 150/100/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, no background noise, normally alert, slightly relaxed, no disfluency
(normal-paced, fairly steady, monologue, formal) 市场整体赚钱效应不佳,还是谨慎行事为好。大家先保住好自己,今年的利润,安稳,迎接春节才是最好的选择。
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 1.1/10; 8.7s, ZH.
ZH_B00009_S06684_W000003 · in -21.5 dBFS · gain +1.5 dB · emolia-03366
(measured, steady, formal, monologue) 距离牛年结束,还有四十天预测的内容,该发生的都发生了,这一年大家也都有所收获。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; average recording, no background noise; genuineness 0.0/6; vocal-burst blend 1.5/10; 6.7s, ZH.
ZH_B00009_S06684_W000004 · in -19.2 dBFS · gain -0.8 dB · emolia-03366
(normal-paced, fairly steady, formal, monologue) 赚到钱的,该揣兜揣兜,该行善的行善。没赚到钱的也不要着急,财不入吉门,放平心态养精蓄锐,静待虎年的到来,明年的报告也会来的。至于谁能把握住机会,就是随缘了。
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 2.0/10; 13.8s, ZH.
ZH_B00009_S06684_W000005 · in -20.2 dBFS · gain +0.2 dB · emolia-03366
(measured, steady, didactic, formal) 孔子曰,无欲速无见小利,欲速则不达,见小利,则大事不成。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 1.3/10; 5.8s, ZH.
ZH_B00009_S06684_W000006 · in -20.1 dBFS · gain +0.1 dB · emolia-03366
Emotional Numbness ↓  /  Concentrationidentity −0.05 emotion 135 %   k-PXR-k4 · #15

This chain comes from the proxy rule: the same two-sided test as above, but because Concentration is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Concentration clearly present — 0.64, higher than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.32.

At the same time Emotional Numbness goes the other way, from 0.92 (higher than 92 % of clips in this corpus) to 0.66 (higher than 66 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.08, then +0.04, then +0.19 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.83 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.83 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.83. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 33 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.758 before conversion and 0.709 after — it fell by 0.049. Neighbour-to-neighbour the worst pair went 0.825 → 0.697. (The earlier render, with segment 1 left raw, scores 0.660 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.318 in the original and +0.428 after conversion — 135 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Emotional Numbness, -0.262 became -0.380.

Quality. Mean predicted overall quality across the segments went 2.62 → 2.82 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.758 → 0.709 -0.049identity cos neighbours 0.825 → 0.697d_b rescored +0.318 → +0.428d_a rescored -0.262 → -0.380d_a mined -0.262d_b mined 0.319min_cos_consec (site) 0.8294min_cos_anchor (site) 0.8294dataset emolialang enspeaker EN_UmybJnfTuqctotal 31.6schain gain +4.3 dBseam step 1.4 dBcrossfades 150/150/100 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a child feminine voice · fairly smooth, slightly relaxed, clear, light breath
(emotional numbness, relief · measured, subdued, steady, whispered) And finally, do not leave cells blank or merge cells.
full caption & clip details
A child feminine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is slightly cool, dark, fairly smooth, thin; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, relief; style: whispered, monologue; average recording, no background noise; genuineness 1.0/6; vocal-burst blend 0.2/10; 4.6s, EN.
EN_UmybJnfTuqc_W000099 · in -19.7 dBFS · gain -0.3 dB · emolia-01972
(normal-paced, normally alert, fairly steady, casual) Tables should be used for organizing data.
full caption & clip details
A child feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; no dominant emotion; style: casual, monologue; good recording, no background noise; genuineness 1.7/6; vocal-burst blend 3.2/10; 3.0s, EN.
EN_UmybJnfTuqc_W000100 · in -16.4 dBFS · gain -3.5 dB · emolia-01972
(normal-paced, normally alert, steady, whispered) Keeping the tables as simple as possible and avoid nesting tables. So tables within tables.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is slightly cool, neutral-bright, fairly smooth, thin; clear, almost no disfluency, fairly narrow pitch, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: whispered, monologue; average recording, quiet background; genuineness 1.0/6; vocal-burst blend 0.0/10; 6.9s, EN.
EN_UmybJnfTuqc_W000101 · in -17.9 dBFS · gain -2.1 dB · emolia-01972
(concentration · normal-paced, normally alert, fairly steady, formal) First, you'll notice that the title is part of the table and it has been merged across multiple cells. Secondly, there are also merged cells in the table. This is a better example of the same information. Notice that the title has been taken out of the table.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as concentration; style: formal, monologue; average recording, quiet background; genuineness 0.0/6; vocal-burst blend 0.7/10; 17.6s, EN.
EN_UmybJnfTuqc_W000103 · in -19.8 dBFS · gain -0.2 dB · emolia-01972
Sourness ↓  /  Emotional Numbnessidentity +0.02 emotion 93 %   k-PXR-k4 · #16

This chain comes from the proxy rule: the same two-sided test as above, but because Emotional Numbness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Emotional Numbness around average — 0.54, higher than 54 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.45.

At the same time Sourness goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.66 (higher than 66 % of clips in this corpus), a change of -0.32. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.10, then +0.17, then +0.18 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.83 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.87 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.83. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 57 s · fr · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.784 before conversion and 0.802 after — it rose by 0.018. Neighbour-to-neighbour the worst pair went 0.856 → 0.760. (The earlier render, with segment 1 left raw, scores 0.709 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.454 in the original and +0.422 after conversion — 93 % of the delta retained, which is essentially all of it. On the other named axis, Sourness, -0.321 became -0.274.

Quality. Mean predicted overall quality across the segments went 3.02 → 3.18 (+0.15) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.784 → 0.802 +0.018identity cos neighbours 0.856 → 0.760d_b rescored +0.454 → +0.422d_a rescored -0.321 → -0.274d_a mined -0.321d_b mined 0.454min_cos_consec (site) 0.8746min_cos_anchor (site) 0.8267dataset emolialang frspeaker FR_JThu7Jf0OeYtotal 56.4schain gain +1.4 dBseam step 0.3 dBcrossfades 150/150/150 ms
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, balanced body, average recording, quiet background, normally alert, slightly relaxed, fairly steady
(sourness, shame, anger · normal-paced, some disfluency, somewhat unclear, monologue) Je vous laisse juger des contenus respectifs des chaînes de ces deux auteurs, et je vous laisse me démontrer que Mademoiselle Wotta est deux fois plus pertinente, ou deux fois plus dissidente, ou deux fois plus cohérente, ou deux fois plus intéressante que, (low mumble) euh, Monsieur Rougeron. Pour le, ah, pour ma part, je n'en suis pas totalement, totalement, (low mumble) euh, convaincu.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as sourness, shame, anger; style: monologue, authoritative; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 3.7/10; 20.0s, FR.
FR_JThu7Jf0OeY_W000071 · in -20.4 dBFS · gain +0.4 dB · emolia-02831
(awe, interest · normal-paced, some disfluency, somewhat unclear, monologue) en histoire, euh, (low mumble) j'avais pour une fois, abondance de, de bien, abondance de choix, donc il paraît que ça ne nuit pas. Donc, nous avons deux femmes contre un homme, les, nous avons Charlie, je sais pas quoi, des Revues du Monde, et (low mumble) euh, l'autre, je sais plus comment elle s'appelle, de, euh, c'est une autre histoire.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as awe, interest; style: monologue, authoritative; average recording, quiet background; genuineness 2.8/6; vocal-burst blend 1.9/10; 16.9s, FR.
FR_JThu7Jf0OeY_W000072 · in -19.0 dBFS · gain -1.0 dB · emolia-02831
(normal-paced, some disfluency, average clarity, monologue) 392 000 abonnés, et pour c'est une autre histoire, 1441 euros mensuels pour 163 900 abonnés, ce qui représente donc 0,2 centimes par abonné pour les revues du monde, 0,9 centimes par abonné
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, casual; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 2.9/10; 13.0s, FR.
FR_JThu7Jf0OeY_W000073 · in -20.8 dBFS · gain +0.8 dB · emolia-02831
(emotional numbness · measured, frequent disfluency, somewhat unclear, didactic) pour, c'est une autre histoire, et seulement pour l'homme, 0,1 centime. Donc, un rapport de fois deux.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: didactic, monologue; average recording, quiet background; genuineness 2.1/6; vocal-burst blend 0.8/10; 7.2s, FR.
FR_JThu7Jf0OeY_W000074 · in -19.8 dBFS · gain -0.2 dB · emolia-02831
Embarrassment ↓  /  Shameidentity −0.05 emotion 89 %   k-PXR-k4 · #17

This chain comes from the proxy rule: the same two-sided test as above, but because Shame is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Shame clearly present — 0.70, higher than 70 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.27.

At the same time Embarrassment goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.71 (higher than 71 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.24, then -0.02, then +0.05 — not a clean run: step 2 moves back the other way by 0.02 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.86 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.90 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.86. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 60 s · en · podcast

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.784 before conversion and 0.732 after — it fell by 0.052. Neighbour-to-neighbour the worst pair went 0.866 → 0.812. (The earlier render, with segment 1 left raw, scores 0.724 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Shame moved +0.269 in the original and +0.240 after conversion — 89 % of the delta retained, which is most of it. On the other named axis, Embarrassment, -0.258 became -0.237.

Quality. Mean predicted overall quality across the segments went 3.01 → 3.16 (+0.15) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.784 → 0.732 -0.052identity cos neighbours 0.866 → 0.812d_b rescored +0.269 → +0.240d_a rescored -0.258 → -0.237d_a mined -0.258d_b mined 0.269min_cos_consec (site) 0.9007min_cos_anchor (site) 0.8650dataset podcastlang enspeaker 530226total 59.3schain gain +3.0 dBseam step 5.8 dBcrossfades 150/150/100 ms
Script — 4 chunks, 3 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normal-paced, normally alert, average clarity, moderate pitch range
(embarrassment · slightly relaxed, fairly steady, frequent disfluency, casual) one that jumps out to me immediately is I was talking to Malcolm Gladwell, (low mumble) uh, best-selling author, podcaster, uh
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as embarrassment; style: casual, conversational; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 3.8/10; 7.8s, EN.
530226_00065592 · in -26.8 dBFS · gain +6.8 dB · podcast-05038
(contemplation, infatuation, shame · relaxed, fairly steady, frequent disfluency, casual) he I had asked him because I (low mumble) uh, you know, as a as a peer of his, I was curious how he filters his projects. How does he think of what he should and shouldn't do? What is a Malcolm Gladwell project to him? And
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as contemplation, infatuation, shame; style: casual, conversational; average recording, no background noise; genuineness 4.7/6; vocal-burst blend 8.7/10; 15.0s, EN.
530226_00066644 · in -27.3 dBFS · gain +7.3 dB · podcast-02992
(contemplation, doubt, affection · neutral tension, moderately variable, some disfluency, casual) because you know, Malcolm Gladwell has a brand and he has a voice and he does things that kind of feel distinctively, Malcolm Gladwell. And I was curious how he defined that for himself. And he said he doesn't do that, he doesn't think of himself uh (low mumble)
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as contemplation, doubt, affection; style: casual, conversational; good recording, quiet background; genuineness 4.5/6; vocal-burst blend 6.6/10; 12.7s, EN.
530226_00068144 · in -26.2 dBFS · gain +6.2 dB · podcast-02973
(shame, infatuation, concentration · slightly relaxed, fairly steady, some disfluency, casual) as as having one particular kind of voice. He doesn't think of himself as a brand. And then he said, he said this line that I immediately jotted down. He said, self-conceptions are powerfully limiting, which is to say that if you have some particular vision of what you are, some particular definition of what you are. Well, then you're going to close off all of these other opportunities that don't fit that narrow window.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as shame, infatuation, concentration; style: casual, monologue; average recording, quiet background; genuineness 1.9/6; vocal-burst blend 5.3/10; 24.2s, EN.
530226_00069416 · in -26.2 dBFS · gain +6.2 dB · podcast-06326
Concentration ↓  /  Prideidentity +0.02 emotion 20 %   k-PXR-k4 · #18

This chain comes from the proxy rule: the same two-sided test as above, but because Pride is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Pride clearly present — 0.67, higher than 67 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.29.

At the same time Concentration goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.67 (higher than 67 % of clips in this corpus), a change of -0.29. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.11, then +0.22, then -0.03 — not a clean run: step 3 moves back the other way by 0.03 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.92 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.90 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.92), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 54 s · no · eurospeech

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.902 before conversion and 0.920 after — it rose by 0.018. Neighbour-to-neighbour the worst pair went 0.893 → 0.925. (The earlier render, with segment 1 left raw, scores 0.745 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Pride moved +0.293 in the original and +0.059 after conversion — 20 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Concentration, -0.290 became -0.340.

Quality. Mean predicted overall quality across the segments went 3.18 → 3.46 (+0.28) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.902 → 0.920 +0.018identity cos neighbours 0.893 → 0.925d_b rescored +0.293 → +0.059d_a rescored -0.290 → -0.340d_a mined -0.290d_b mined 0.293min_cos_consec (site) 0.8970min_cos_anchor (site) 0.9210dataset eurospeechlang nospeaker norway_9531-1total 52.5schain gain +2.0 dBseam step 0.8 dBcrossfades 150/150/150 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a middle-aged masculine voice · neutral-bright, balanced body, quiet background
(concentration, bitterness, sadness · measured, energised, neutral tension, cartoonish) Landet har stort sett vært i konflikt og i en krigslignende tilstand. Det har få venner i omverdenen. Jeg mener det er viktig at Norge deltar aktivt i internasjonalt samarbeid for å hindre en total statskollaps i Eritrea,
full caption & clip details
A middle-aged masculine voice; delivery is energised, measured, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, wide pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, bitterness, sadness; style: cartoonish, storytelling; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 3.7/10; 15.0s, NO.
norway_9531-1_17121424_17136384 · in -32.5 dBFS · gain +12.5 dB · eurospeech-02334
(disappointment, impatience and irritability, concentration · normal-paced, normally alert, slightly relaxed, monologue) med de følger det kan få for en allerede hardt prøvet befolkning, og den (ahem) destabilisering det kan bety for den konfliktrammede regionen rundt Afrikas Horn. Det er ikke naturlig med noe omfattende tosidig utviklingssamarbeid med et sånt regime,
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disappointment, impatience and irritability, concentration; style: monologue, storytelling; good recording, quiet background; genuineness 1.5/6; vocal-burst blend 4.9/10; 14.7s, NO.
norway_9531-1_17136384_17151088 · in -31.9 dBFS · gain +11.9 dB · eurospeech-02334
(pride, shame, triumph · measured, normally alert, slightly relaxed, narration) men når mennesker er i desperat behov for nødhjelp, får vi og det internasjonale samfunn behov for å stille opp. Så får vi støtte politiske prosesser for å bringe utviklinga inn på en bedre vei.
full caption & clip details
A child masculine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; slurred, some disfluency, wide pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; reads as pride, shame, triumph; style: narration, storytelling; average recording, quiet background; genuineness 1.7/6; vocal-burst blend 4.1/10; 12.9s, NO.
norway_9531-1_17151088_17164032 · in -32.1 dBFS · gain +12.1 dB · eurospeech-02334
(pride · normal-paced, normally alert, neutral tension, dramatic) en bedre vei. Jeg mener situasjonen i Eritrea også mye viser noen av de dilemmaene vi har i internasjonal politikk. Menneskerettigheter brytes i en rekke land,
full caption & clip details
A child masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, wide pitch range, audible breath; affect is positive, neutral stance, slightly guarded; reads as pride; style: dramatic, monologue; average recording, quiet background; genuineness 2.0/6; vocal-burst blend 1.5/10; 10.5s, NO.
norway_9531-1_17164032_17174528 · in -31.6 dBFS · gain +11.6 dB · eurospeech-02334
Emotional Numbness ↓  /  Concentrationidentity −0.02 emotion 95 %   k-PXR-k4 · #19

This chain comes from the proxy rule: the same two-sided test as above, but because Concentration is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Concentration clearly present — 0.68, higher than 68 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.29.

At the same time Emotional Numbness goes the other way, from 0.93 (higher than 93 % of clips in this corpus) to 0.44 (lower than 56 % of clips in this corpus), a change of -0.49. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are -0.04, then +0.22, then +0.11 — not a clean run: step 1 moves back the other way by 0.04 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.91 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.90 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.91), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 30 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.773 before conversion and 0.752 after — it fell by 0.021. Neighbour-to-neighbour the worst pair went 0.751 → 0.782. (The earlier render, with segment 1 left raw, scores 0.721 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.291 in the original and +0.276 after conversion — 95 % of the delta retained, which is essentially all of it. On the other named axis, Emotional Numbness, -0.489 became -0.152.

Quality. Mean predicted overall quality across the segments went 2.57 → 2.87 (+0.30) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.773 → 0.752 -0.021identity cos neighbours 0.751 → 0.782d_b rescored +0.291 → +0.276d_a rescored -0.489 → -0.152d_a mined -0.489d_b mined 0.291min_cos_consec (site) 0.8962min_cos_anchor (site) 0.9108dataset emolialang enspeaker EN_Cf4SW2WKIVktotal 29.1schain gain +2.3 dBseam step 1.1 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · slightly cool, neutral-bright, rough, highly aroused, almost no disfluency, wide pitch range
(emotional numbness · normal-paced, slightly relaxed, fairly steady, authoritative) Unified system of separation, retirement and pension.
full caption & clip details
An adult masculine voice; delivery is highly aroused, normal-paced, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, rough, thin; clear, almost no disfluency, wide pitch range, light breath; affect is neutral, dominant, fairly guarded; reads as emotional numbness; style: authoritative, formal; good recording, quiet background; genuineness 0.7/6; vocal-burst blend 0.0/10; 3.5s, EN.
EN_Cf4SW2WKIVk_W000449 · in -20.2 dBFS · gain +0.2 dB · emolia-00567
(pride · normal-paced, tense, moderately variable, authoritative) This grants a monthly disability pension in Liu.
full caption & clip details
An adult masculine voice; delivery is highly aroused, normal-paced, tense, moderately variable; timbre is slightly cool, neutral-bright, rough, thin; very clear, almost no disfluency, wide pitch range, audible breath; affect is neutral, dominant, fairly guarded; reads as pride; style: authoritative, dramatic; below-average recording, quiet background; genuineness 1.2/6; vocal-burst blend 0.1/10; 3.5s, EN.
EN_Cf4SW2WKIVk_W000450 · in -18.5 dBFS · gain -1.5 dB · emolia-00567
(malevolence malice · measured, slightly relaxed, fairly steady, authoritative) It promotes the use of internet, internet, and other ICT to provide opportunities for citizens.
full caption & clip details
A middle-aged masculine voice; delivery is highly aroused, measured, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, rough, balanced body; very clear, almost no disfluency, wide pitch range, normal breath; affect is neutral, dominant, fairly guarded; reads as malevolence malice; style: authoritative, dramatic; below-average recording, quiet background; genuineness 0.5/6; vocal-burst blend 0.0/10; 11.0s, EN.
EN_Cf4SW2WKIVk_W000452 · in -21.6 dBFS · gain +1.6 dB · emolia-00567
(concentration, malevolence malice · brisk, neutral tension, moderately variable, authoritative) This will provide for a rational and holistic management and development of our country's land and water resources. Hold owners accountable for making these lands productive and sustainable.
full caption & clip details
A middle-aged masculine voice; delivery is highly aroused, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, rough, thin; very clear, almost no disfluency, wide pitch range, normal breath; affect is neutral, dominant, fairly guarded; reads as concentration, malevolence malice; style: authoritative, dramatic; below-average recording, some background noise; genuineness 0.3/6; vocal-burst blend 0.7/10; 11.7s, EN.
EN_Cf4SW2WKIVk_W000453 · in -20.3 dBFS · gain +0.3 dB · emolia-00567
Relief ↓  /  Emotional Numbnessidentity −0.08 emotion 100 %   k-PXR-k4 · #20

This chain comes from the proxy rule: the same two-sided test as above, but because Emotional Numbness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Emotional Numbness around average — 0.58, higher than 58 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.41.

At the same time Relief goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.52 (higher than 52 % of clips in this corpus), a change of -0.47. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.13, then +0.05, then +0.23 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.90 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.90 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.90), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 62 s · en · emolia

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.783 before conversion and 0.699 after — it fell by 0.084. Neighbour-to-neighbour the worst pair went 0.783 → 0.699. (The earlier render, with segment 1 left raw, scores 0.699 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.414 in the original and +0.414 after conversion — 100 % of the delta retained, which is essentially all of it. On the other named axis, Relief, -0.469 became -0.537.

Quality. Mean predicted overall quality across the segments went 2.95 → 3.15 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.783 → 0.699 -0.084identity cos neighbours 0.783 → 0.699d_b rescored +0.414 → +0.414d_a rescored -0.469 → -0.537d_a mined -0.469d_b mined 0.414min_cos_consec (site) 0.9007min_cos_anchor (site) 0.9007dataset emolialang enspeaker EN_QgwpVXZ5liutotal 61.1schain gain +2.3 dBseam step 2.0 dBcrossfades 150/150/150 ms
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · neutral-toned, slightly dark, balanced body, average recording, quiet background, measured, slightly relaxed, steady
(relief · subdued, light breath, didactic, monologue) And this delta positive charge remains here and so one alcohol group will leave.
full caption & clip details
A young adult masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as relief; style: didactic, monologue; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 0.5/10; 6.7s, EN.
EN_QgwpVXZ5liu_W000203 · in -19.2 dBFS · gain -0.8 dB · emolia-00832
(triumph, sexual lust, infatuation · subdued, light breath, monologue, didactic) Creating one new SiOH linkage. This acid catalyzed reactions normally take place at pH of less than 2.2 and it has a fast protonation step and the silicon becomes electrophilic, uh, (low mumble) after the protonation and therefore, is more susceptible to attack by water and the protonation becomes slower
full caption & clip details
A young adult masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as triumph, sexual lust, infatuation; style: monologue, didactic; average recording, quiet background; genuineness 1.8/6; vocal-burst blend 1.4/10; 27.6s, EN.
EN_QgwpVXZ5liu_W000204 · in -23.3 dBFS · gain +3.3 dB · emolia-00832
(concentration · subdued, normal breath, didactic, monologue) Now in basic conditions the OH minus group attacks the silane, (low mumble) uh, tetra alkoxide silane and you get this kind of a (low mumble) OH delta minus charge here and
full caption & clip details
A young adult masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: didactic, monologue; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 0.5/10; 16.5s, EN.
EN_QgwpVXZ5liu_W000206 · in -22.0 dBFS · gain +2.0 dB · emolia-00832
(emotional numbness, pain · very low-energy, normal breath, didactic, monologue) So, this also gets a OR delta minus chart and then this OR minus leaves and you are left with a new
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, pain; style: didactic, monologue; average recording, quiet background; genuineness 2.1/6; vocal-burst blend 0.0/10; 11.0s, EN.
EN_QgwpVXZ5liu_W000207 · in -21.9 dBFS · gain +1.9 dB · emolia-00832