The 10 scarcest languages, 2 chains each (supply 1,685-11,067).
Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
The chain starts with style: ranting (S_RANT) around average — 0.47, lower than 53 % of clips in this corpus — and ends with it above average at 0.72, higher than 72 % of clips in this corpus. That is a total rise of 0.25.
It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.92 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.92 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.92), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.932 before conversion and 0.840 after — it fell by 0.092. Neighbour-to-neighbour the worst pair went 0.932 → 0.840. (The earlier render, with segment 1 left raw, scores 0.848 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, S_RANT — style: ranting moved +0.239 in the original and +0.289 after conversion — 121 % of the delta retained, i.e. the move came out slightly larger after conversion than before.
Quality. Mean predicted overall quality across the segments went 3.00 → 3.22 (+0.22) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
The chain starts with tension (TENS) high — 0.85, higher than 85 % of clips in this corpus — and works its way down to above average at 0.63, higher than 63 % of clips in this corpus. That is a total fall of 0.23.
It takes 2 clips to get there. Clip to clip the moves are -0.23 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.44 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.44 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.44, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.452 before conversion and 0.789 after — it rose by 0.337. Neighbour-to-neighbour the worst pair went 0.452 → 0.789. (The earlier render, with segment 1 left raw, scores 0.765 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, TENS — tension moved -0.234 in the original and -0.643 after conversion — 275 % of the delta retained, i.e. the move came out slightly larger after conversion than before.
Quality. Mean predicted overall quality across the segments went 3.05 → 3.42 (+0.38) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
The chain starts with style: newsreading (S_NEWS) below average — 0.36, lower than 64 % of clips in this corpus — and works its way down to low at 0.11, lower than 89 % of clips in this corpus. That is a total fall of 0.25.
It takes 2 clips to get there. Clip to clip the moves are -0.25 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.72 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.72 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.72, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.699 before conversion and 0.756 after — it rose by 0.056. Neighbour-to-neighbour the worst pair went 0.699 → 0.756. (The earlier render, with segment 1 left raw, scores 0.567 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, S_NEWS — style: newsreading moved -0.248 in the original and -0.101 after conversion — 41 % of the delta retained, so a meaningful part of the trajectory was flattened.
Quality. Mean predicted overall quality across the segments went 2.95 → 3.06 (+0.10) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
The chain starts with resonance: oral (R_ORAL) low — 0.17, lower than 83 % of clips in this corpus — and ends with it below average at 0.40, lower than 60 % of clips in this corpus. That is a total rise of 0.23.
It takes 2 clips to get there. Clip to clip the moves are +0.23 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.93 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.93 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.93), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.932 before conversion and 0.922 after — it fell by 0.009. Neighbour-to-neighbour the worst pair went 0.932 → 0.922. (The earlier render, with segment 1 left raw, scores 0.817 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, R_ORAL — resonance: oral moved +0.227 in the original and +0.149 after conversion — 66 % of the delta retained.
Quality. Mean predicted overall quality across the segments went 2.65 → 3.13 (+0.48) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
The chain starts with style: casual (S_CASU) at the very bottom of the range — 0.06, lower than 94 % of clips in this corpus — and ends with it below average at 0.30, lower than 70 % of clips in this corpus. That is a total rise of 0.24.
It takes 2 clips to get there. Clip to clip the moves are +0.24 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.88 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.88 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.88. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.831 before conversion and 0.850 after — it rose by 0.019. Neighbour-to-neighbour the worst pair went 0.831 → 0.850. (The earlier render, with segment 1 left raw, scores 0.730 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, S_CASU — style: casual moved +0.241 in the original and +0.123 after conversion — 51 % of the delta retained.
Quality. Mean predicted overall quality across the segments went 3.02 → 3.38 (+0.35) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
The chain starts with style: authoritative (S_AUTH) around average — 0.43, lower than 57 % of clips in this corpus — and works its way down to low at 0.22, lower than 78 % of clips in this corpus. That is a total fall of 0.21.
It takes 2 clips to get there. Clip to clip the moves are -0.21 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.93 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.93 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.93), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.925 before conversion and 0.921 after — it fell by 0.004. Neighbour-to-neighbour the worst pair went 0.925 → 0.921. (The earlier render, with segment 1 left raw, scores 0.760 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, S_AUTH — style: authoritative moved -0.206 in the original and -0.187 after conversion — 91 % of the delta retained, which is essentially all of it.
Quality. Mean predicted overall quality across the segments went 3.10 → 3.31 (+0.21) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
The chain starts with perceived age (AGEV) around average — 0.49, right about the corpus median — and ends with it above average at 0.74, higher than 74 % of clips in this corpus. That is a total rise of 0.25.
It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.29 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.29 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.29, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.297 before conversion and 0.824 after — it rose by 0.527. Neighbour-to-neighbour the worst pair went 0.297 → 0.824. (The earlier render, with segment 1 left raw, scores 0.672 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
The emotional move did not survive. Re-scored end to end, AGEV — perceived age moved +0.247 in the original and -0.008 after conversion — it changed direction. On this chain the corrected audio is not an improvement.
Quality. Mean predicted overall quality across the segments went 2.90 → 3.29 (+0.39) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
The chain starts with perceived age (AGEV) below average — 0.33, lower than 67 % of clips in this corpus — and ends with it around average at 0.58, higher than 58 % of clips in this corpus. That is a total rise of 0.25.
It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.93 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.93 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.93), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.912 before conversion and 0.951 after — it rose by 0.039. Neighbour-to-neighbour the worst pair went 0.912 → 0.951. (The earlier render, with segment 1 left raw, scores 0.663 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
The emotional move did not survive. Re-scored end to end, AGEV — perceived age moved +0.255 in the original and -0.021 after conversion — it changed direction. On this chain the corrected audio is not an improvement.
Quality. Mean predicted overall quality across the segments went 2.94 → 3.27 (+0.33) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
The chain starts with emphasis / stress strength (EMPH) at the very top of the range — 0.94, higher than 94 % of clips in this corpus — and works its way down to above average at 0.69, higher than 69 % of clips in this corpus. That is a total fall of 0.25.
It takes 2 clips to get there. Clip to clip the moves are -0.25 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores -0.02 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst -0.02 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (-0.02, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.037 before conversion and 0.787 after — it rose by 0.751. Neighbour-to-neighbour the worst pair went 0.037 → 0.787. (The earlier render, with segment 1 left raw, scores 0.465 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, EMPH — emphasis / stress strength moved -0.257 in the original and -0.026 after conversion — 10 % of the delta retained, so a meaningful part of the trajectory was flattened.
Quality. Mean predicted overall quality across the segments went 3.02 → 3.25 (+0.23) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
The chain starts with style: cartoonish (S_CART) low — 0.16, lower than 84 % of clips in this corpus — and ends with it below average at 0.41, lower than 59 % of clips in this corpus. That is a total rise of 0.25.
It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.83 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.83 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.83. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.754 before conversion and 0.822 after — it rose by 0.068. Neighbour-to-neighbour the worst pair went 0.754 → 0.822. (The earlier render, with segment 1 left raw, scores 0.451 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, S_CART — style: cartoonish moved +0.252 in the original and +0.069 after conversion — 27 % of the delta retained, so a meaningful part of the trajectory was flattened.
Quality. Mean predicted overall quality across the segments went 2.74 → 3.31 (+0.58) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
The chain starts with metallic quality (METL) low — 0.21, lower than 79 % of clips in this corpus — and ends with it around average at 0.45, lower than 55 % of clips in this corpus. That is a total rise of 0.24.
It takes 2 clips to get there. Clip to clip the moves are +0.24 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.96 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.96 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.96), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.952 before conversion and 0.957 after — it rose by 0.006. Neighbour-to-neighbour the worst pair went 0.952 → 0.957. (The earlier render, with segment 1 left raw, scores 0.856 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, METL — metallic quality moved +0.257 in the original and +0.223 after conversion — 87 % of the delta retained, which is most of it.
Quality. Mean predicted overall quality across the segments went 3.33 → 3.61 (+0.28) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
The chain starts with stance / assertiveness (STNC) below average — 0.33, lower than 67 % of clips in this corpus — and ends with it above average at 0.58, higher than 58 % of clips in this corpus. That is a total rise of 0.25.
It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.53 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.53 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.53, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.502 before conversion and 0.642 after — it rose by 0.140. Neighbour-to-neighbour the worst pair went 0.502 → 0.642. (The earlier render, with segment 1 left raw, scores 0.515 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
The emotional move did not survive. Re-scored end to end, STNC — stance / assertiveness moved +0.232 in the original and -0.056 after conversion — it changed direction. On this chain the corrected audio is not an improvement.
Quality. Mean predicted overall quality across the segments went 2.18 → 3.26 (+1.08) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
The chain starts with harshness of articulation (ARSH) above average — 0.58, higher than 58 % of clips in this corpus — and works its way down to below average at 0.33, lower than 67 % of clips in this corpus. That is a total fall of 0.25.
It takes 2 clips to get there. Clip to clip the moves are -0.25 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the eurospeech clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.951 before conversion and 0.939 after — it fell by 0.011. Neighbour-to-neighbour the worst pair went 0.951 → 0.939. (The earlier render, with segment 1 left raw, scores 0.835 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, ARSH — harshness of articulation moved -0.249 in the original and -0.063 after conversion — 25 % of the delta retained, so a meaningful part of the trajectory was flattened.
Quality. Mean predicted overall quality across the segments went 2.97 → 3.41 (+0.44) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
The chain starts with vocal focus (FOCS) high — 0.80, higher than 80 % of clips in this corpus — and works its way down to around average at 0.55, higher than 55 % of clips in this corpus. That is a total fall of 0.25.
It takes 2 clips to get there. Clip to clip the moves are -0.25 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the eurospeech clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.927 before conversion and 0.946 after — it rose by 0.019. Neighbour-to-neighbour the worst pair went 0.927 → 0.946. (The earlier render, with segment 1 left raw, scores 0.862 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, FOCS — vocal focus moved -0.246 in the original and -0.229 after conversion — 93 % of the delta retained, which is essentially all of it.
Quality. Mean predicted overall quality across the segments went 2.97 → 3.26 (+0.29) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
The chain starts with resonance: mixed (R_MIXD) above average — 0.69, higher than 69 % of clips in this corpus — and works its way down to around average at 0.47, lower than 53 % of clips in this corpus. That is a total fall of 0.22.
It takes 2 clips to get there. Clip to clip the moves are -0.22 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the eurospeech clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.930 before conversion and 0.868 after — it fell by 0.062. Neighbour-to-neighbour the worst pair went 0.930 → 0.868. (The earlier render, with segment 1 left raw, scores 0.674 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, R_MIXD — resonance: mixed moved -0.223 in the original and -0.156 after conversion — 70 % of the delta retained, which is most of it.
Quality. Mean predicted overall quality across the segments went 3.12 → 3.32 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
The chain starts with tension (TENS) around average — 0.51, higher than 51 % of clips in this corpus — and ends with it high at 0.76, higher than 76 % of clips in this corpus. That is a total rise of 0.25.
It takes 2 clips to get there. Clip to clip the moves are +0.25 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the eurospeech clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.912 before conversion and 0.920 after — it rose by 0.008. Neighbour-to-neighbour the worst pair went 0.912 → 0.920. (The earlier render, with segment 1 left raw, scores 0.844 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, TENS — tension moved +0.246 in the original and +0.251 after conversion — 102 % of the delta retained, which is essentially all of it.
Quality. Mean predicted overall quality across the segments went 3.03 → 3.30 (+0.27) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
The chain starts with perceived gender (GEND) around average — 0.48, lower than 52 % of clips in this corpus — and ends with it above average at 0.69, higher than 69 % of clips in this corpus. That is a total rise of 0.22.
It takes 2 clips to get there. Clip to clip the moves are +0.22 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.908 before conversion and 0.911 after — it rose by 0.004. Neighbour-to-neighbour the worst pair went 0.908 → 0.911. (The earlier render, with segment 1 left raw, scores 0.728 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, GEND — perceived gender moved +0.219 in the original and +0.200 after conversion — 91 % of the delta retained, which is essentially all of it.
Quality. Mean predicted overall quality across the segments went 3.07 → 3.35 (+0.27) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
The chain starts with chunking / phrasing density (CHNK) at the very top of the range — 0.98, higher than 98 % of clips in this corpus — and works its way down to above average at 0.74, higher than 74 % of clips in this corpus. That is a total fall of 0.24.
It takes 2 clips to get there. Clip to clip the moves are -0.24 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? No similarity score is available here — the mls clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.928 before conversion and 0.928 after — it fell by 0.001. Neighbour-to-neighbour the worst pair went 0.928 → 0.928. (The earlier render, with segment 1 left raw, scores 0.798 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, CHNK — chunking / phrasing density moved -0.241 in the original and -0.218 after conversion — 90 % of the delta retained, which is essentially all of it.
Quality. Mean predicted overall quality across the segments went 3.10 → 3.35 (+0.25) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
The chain starts with vocal focus (FOCS) around average — 0.54, higher than 54 % of clips in this corpus — and works its way down to below average at 0.29, lower than 71 % of clips in this corpus. That is a total fall of 0.24.
It takes 2 clips to get there. Clip to clip the moves are -0.24 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.72 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.72 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.72, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.693 before conversion and 0.808 after — it rose by 0.115. Neighbour-to-neighbour the worst pair went 0.693 → 0.808. (The earlier render, with segment 1 left raw, scores 0.698 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, FOCS — vocal focus moved -0.252 in the original and -0.638 after conversion — 253 % of the delta retained, i.e. the move came out slightly larger after conversion than before.
Quality. Mean predicted overall quality across the segments went 2.80 → 3.00 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
The chain starts with roughness (ROUG) high — 0.86, higher than 86 % of clips in this corpus — and works its way down to above average at 0.62, higher than 62 % of clips in this corpus. That is a total fall of 0.25.
It takes 2 clips to get there. Clip to clip the moves are -0.25 — a single step, so there is no internal shape to speak of.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.90 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.90 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.90), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.904 before conversion and 0.852 after — it fell by 0.051. Neighbour-to-neighbour the worst pair went 0.904 → 0.852. (The earlier render, with segment 1 left raw, scores 0.769 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, ROUG — roughness moved -0.253 in the original and -0.255 after conversion — 101 % of the delta retained, which is essentially all of it.
Quality. Mean predicted overall quality across the segments went 2.83 → 3.27 (+0.44) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.