proxy_taillift__PXR__T0.70__C0.25__INTERNAL — voice-corrected

Manifest tier. proxy_taillift, rule PXR, T=0.7, step cap 0.25. Population 2 chains (0 h) over 6 corpora. The SHAREABLE variant of this tier (podcast and evasnippets excluded) holds 2.

This is the voice-conversion-corrected version of this tier. The original, uncorrected page is still there and unchanged: t_proxy_taillift__PXR__T0.70__C0.25__INTERNAL.html. Every card below carries both renders so you can switch between them without leaving the page. What was done, and what it measured →
2chains converted
8segments re-voiced
0.441 → 0.595median worst-to-anchor identity cosine
99 %median emotion delta retained
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken.

(underlined, plain · delivery, style) — the tag before the words. Emotions first, then how it is delivered. Underlined descriptors are the ones that change across this chain — anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words are a real non-speech sound, printed where it happens.

The full generated caption for any clip is under “full caption & clip details”. Its perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.

The tags and captions describe the ORIGINAL clips — they are the corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the converted audio are the ones printed in each card's own paragraph and chips.
Confusion ↓  /  Concentrationidentity +0.36 emotion 115 %   proxy_taillift__PXR__T0.70__C0.25__INTERNAL · #1

This chain comes from the proxy rule: the same two-sided test as above, but because Concentration is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Concentration below average — 0.27, lower than 73 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.71.

At the same time Confusion goes the other way, from 0.73 (higher than 73 % of clips in this corpus) to 0.02 (lower than 98 % of clips in this corpus), a change of -0.70. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.11, then +0.19, then +0.24, then +0.18 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores -0.05 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.02 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (-0.05, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 38 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.057 before conversion and 0.415 after — it rose by 0.358. Neighbour-to-neighbour the worst pair went 0.062 → 0.373. (The earlier render, with segment 1 left raw, scores 0.352 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.711 in the original and +0.817 after conversion — 115 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Confusion, -0.705 became -0.784.

Quality. Mean predicted overall quality across the segments went 2.42 → 2.76 (+0.34) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.057 → 0.415 +0.358identity cos neighbours 0.062 → 0.373d_b rescored +0.711 → +0.817d_a rescored -0.705 → -0.784d_a mined -0.705d_b mined 0.711min_cos_consec (site) 0.0215min_cos_anchor (site) -0.0465dataset emolialang enspeaker EN_8n0eK_7khaktotal 36.8schain gain +3.3 dBseam step 2.8 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · fairly steady
(measured, subdued, relaxed, casual) 250, (low mumble) um, with the (ahem) technical writing and, (low mumble) um.
full caption & clip details
A young adult masculine voice; delivery is subdued, measured, relaxed, fairly steady; timbre is neutral-toned, dark, fairly smooth, slightly thin; slurred, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, submissive, neutral openness; no dominant emotion; style: casual, conversational; below-average recording, quiet background; genuineness 4.2/6; vocal-burst blend 1.9/10; 5.1s, EN.
EN_8n0eK_7khak_W000058 · in -16.4 dBFS · gain -3.6 dB · emolia-00318
(helplessness, sadness · normal-paced, normally alert, slightly relaxed, casual) Wordsmithing and so forth. We feel that for that level.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as helplessness, sadness; style: casual, monologue; average recording, no background noise; genuineness 3.2/6; vocal-burst blend 2.1/10; 3.6s, EN.
EN_8n0eK_7khak_W000059 · in -13.9 dBFS · gain -6.1 dB · emolia-00318
(normal-paced, normally alert, slightly relaxed, monologue) And then we would be further directing them to prepare an RFP.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: monologue, casual; good recording, no background noise; genuineness 2.1/6; vocal-burst blend 2.8/10; 4.4s, EN.
EN_8n0eK_7khak_W000060 · in -16.4 dBFS · gain -3.6 dB · emolia-00318
(normal-paced, normally alert, slightly relaxed, monologue) that would outline the activities that the district needs (ahem) public outreach for.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: monologue, casual; good recording, no background noise; genuineness 2.3/6; vocal-burst blend 1.2/10; 5.7s, EN.
EN_8n0eK_7khak_W000061 · in -16.2 dBFS · gain -3.8 dB · emolia-00318
(concentration · slow, very low-energy, relaxed, monologue) Okay. So you're asking the board for two different things, I think. (low mumble) Um, one, do we increase the budget from 35,000 to 50,000 to, (low mumble) um, are we, (low mumble) uh, approving the district
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is slightly cool, slightly dark, rough, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, submissive, slightly guarded; reads as concentration; style: monologue, didactic; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 0.1/10; 18.9s, EN.
EN_8n0eK_7khak_W000062 · in -19.2 dBFS · gain -0.8 dB · emolia-00318
Fear ↓  /  Astonishment Surpriseidentity −0.05 emotion 83 %   proxy_taillift__PXR__T0.70__C0.25__INTERNAL · #2

This chain comes from the proxy rule: the same two-sided test as above, but because Astonishment Surprise is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.

The chain starts with Astonishment Surprise barely there — 0.14, lower than 86 % of clips in this corpus — and ends with it strongly present at 0.87, higher than 87 % of clips in this corpus. That is a total rise of 0.73.

At the same time Fear goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.25 (lower than 75 % of clips in this corpus), a change of -0.71. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.22, then +0.16, then +0.13, then +0.21 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.93 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.92 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.93), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 47 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.825 before conversion and 0.775 after — it fell by 0.050. Neighbour-to-neighbour the worst pair went 0.672 → 0.540. (The earlier render, with segment 1 left raw, scores 0.617 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Astonishment Surprise moved +0.734 in the original and +0.609 after conversion — 83 % of the delta retained, which is most of it. On the other named axis, Fear, -0.713 became -0.627.

Quality. Mean predicted overall quality across the segments went 2.80 → 2.95 (+0.15) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.825 → 0.775 -0.050identity cos neighbours 0.672 → 0.540d_b rescored +0.734 → +0.609d_a rescored -0.713 → -0.627d_a mined -0.713d_b mined 0.734min_cos_consec (site) 0.9157min_cos_anchor (site) 0.9265dataset emolialang enspeaker EN_c5BYOO0j3Fytotal 45.9schain gain +1.2 dBseam step 1.5 dBcrossfades 150/100/150/100 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(fear, concentration · steady, almost no disfluency, formal, newsreading) This means that the bright hemisphere is visible from Earth when Iapetus is on the western side of Saturn, and that the dark hemisphere is visible when Iapetus is on the eastern side.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fear, concentration; style: formal, newsreading; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.0/10; 9.2s, EN.
EN_c5BYOO0j3Fy_W000008 · in -14.9 dBFS · gain -5.1 dB · emolia-02588
(fairly steady, no disfluency, formal, monologue) The Dark Hemisphere was later named Cassini Regio in his honor
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 2.0/10; 3.5s, EN.
EN_c5BYOO0j3Fy_W000009 · in -13.4 dBFS · gain -6.6 dB · emolia-02588
(fairly steady, no disfluency, formal, monologue) Iopetus is named after the Titan Iopetus from Greek mythology
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 2.2/10; 3.6s, EN.
EN_c5BYOO0j3Fy_W000011 · in -14.4 dBFS · gain -5.6 dB · emolia-02588
(emotional numbness · fairly steady, no disfluency, newsreading, formal) The name was suggested by John Herschel – son of William Herschel, discoverer of Mimas and Enceladus – in his 1847 publication Results of Astronomical Observations Made at the Cape of Good Hope, in which he advocated naming the moons of Saturn after the Titans, brothers and sisters of the Titan Cronus – whom the Romans equated with their god Saturn.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 19.0s, EN.
EN_c5BYOO0j3Fy_W000012 · in -14.8 dBFS · gain -5.2 dB · emolia-02588
(fairly steady, no disfluency, authoritative, formal) When first discovered, Iapetus was among four Saturnian moons labeled the Sedera Lodoisia by their discoverer Giovanni Cassini after King Louis XIV. The other three were Tethys, Dione and Rhea
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: authoritative, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 11.2s, EN.
EN_c5BYOO0j3Fy_W000013 · in -13.7 dBFS · gain -6.3 dB · emolia-02588