AB2 at chain length k=3, all corpora, at the mining floor.
Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
This chain comes from the two-sided rule: it only counts if both emotions move — Fatigue Exhaustion down and Infatuation up — by at least 0.25 each.
The chain starts with Infatuation around average — 0.54, higher than 54 % of clips in this corpus — and ends with it strongly present at 0.85, higher than 85 % of clips in this corpus. That is a total rise of 0.31.
At the same time Fatigue Exhaustion goes the other way, from 0.87 (higher than 87 % of clips in this corpus) to 0.57 (higher than 57 % of clips in this corpus), a change of -0.30. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.09, then +0.23 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.89 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.88 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.89. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.891 before conversion and 0.864 after — it fell by 0.027. Neighbour-to-neighbour the worst pair went 0.871 → 0.855. (The earlier render, with segment 1 left raw, scores 0.816 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Infatuation moved +0.314 in the original and +0.405 after conversion — 129 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Fatigue Exhaustion, -0.296 became -0.318.
Quality. Mean predicted overall quality across the segments went 3.05 → 3.18 (+0.13) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This chain comes from the two-sided rule: it only counts if both emotions move — Malevolence Malice down and Intoxication Altered States of Consciousness up — by at least 0.25 each.
The chain starts with Intoxication Altered States of Consciousness clearly present — 0.58, higher than 58 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.41.
At the same time Malevolence Malice goes the other way, from 0.72 (higher than 72 % of clips in this corpus) to 0.47 (lower than 53 % of clips in this corpus), a change of -0.25. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.23, then +0.17 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.74 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.75 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.74, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.695 before conversion and 0.795 after — it rose by 0.100. Neighbour-to-neighbour the worst pair went 0.751 → 0.798. (The earlier render, with segment 1 left raw, scores 0.647 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Intoxication Altered States of Consciousness moved +0.407 in the original and +0.556 after conversion — 137 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Malevolence Malice, -0.253 became -0.172.
Quality. Mean predicted overall quality across the segments went 2.90 → 3.02 (+0.12) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This chain comes from the two-sided rule: it only counts if both emotions move — Concentration down and Pain up — by at least 0.25 each.
The chain starts with Pain around average — 0.54, higher than 54 % of clips in this corpus — and ends with it strongly present at 0.86, higher than 86 % of clips in this corpus. That is a total rise of 0.32.
At the same time Concentration goes the other way, from 0.90 (higher than 90 % of clips in this corpus) to 0.42 (lower than 58 % of clips in this corpus), a change of -0.48. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.17, then +0.14 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.78 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.73 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.78, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.690 before conversion and 0.616 after — it fell by 0.074. Neighbour-to-neighbour the worst pair went 0.693 → 0.715. (The earlier render, with segment 1 left raw, scores 0.519 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Pain moved +0.317 in the original and +0.305 after conversion — 96 % of the delta retained, which is essentially all of it. On the other named axis, Concentration, -0.477 became -0.583.
Quality. Mean predicted overall quality across the segments went 2.29 → 2.81 (+0.52) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This chain comes from the two-sided rule: it only counts if both emotions move — Impatience and Irritability down and Hope Enthusiasm Optimism up — by at least 0.25 each.
The chain starts with Hope Enthusiasm Optimism clearly present — 0.66, higher than 66 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.33.
At the same time Impatience and Irritability goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.61 (higher than 61 % of clips in this corpus), a change of -0.36. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.11, then +0.22 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.49 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.46 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.49, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.495 before conversion and 0.816 after — it rose by 0.321. Neighbour-to-neighbour the worst pair went 0.439 → 0.741. (The earlier render, with segment 1 left raw, scores 0.654 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Hope Enthusiasm Optimism moved +0.319 in the original and +0.030 after conversion — 9 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Impatience and Irritability, -0.358 became -0.209.
Quality. Mean predicted overall quality across the segments went 2.69 → 3.17 (+0.48) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This chain comes from the two-sided rule: it only counts if both emotions move — Impatience and Irritability down and Jealousy and Envy up — by at least 0.25 each.
The chain starts with Jealousy and Envy around average — 0.53, higher than 53 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.41.
At the same time Impatience and Irritability goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.70 (higher than 70 % of clips in this corpus), a change of -0.27. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.19, then +0.22 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.81 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.84 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.81. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.800 before conversion and 0.853 after — it rose by 0.053. Neighbour-to-neighbour the worst pair went 0.837 → 0.869. (The earlier render, with segment 1 left raw, scores 0.572 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Jealousy and Envy moved +0.386 in the original and +0.419 after conversion — 108 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Impatience and Irritability, -0.278 became -0.334.
Quality. Mean predicted overall quality across the segments went 2.65 → 3.18 (+0.53) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This chain comes from the two-sided rule: it only counts if both emotions move — Contemplation down and Relief up — by at least 0.25 each.
The chain starts with Relief clearly present — 0.66, higher than 66 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.29.
At the same time Contemplation goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.72 (higher than 72 % of clips in this corpus), a change of -0.27. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.15, then +0.14 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.91 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.92 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.91), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.884 before conversion and 0.808 after — it fell by 0.076. Neighbour-to-neighbour the worst pair went 0.866 → 0.795. (The earlier render, with segment 1 left raw, scores 0.747 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Relief moved +0.292 in the original and +0.171 after conversion — 59 % of the delta retained. On the other named axis, Contemplation, -0.269 became -0.229.
Quality. Mean predicted overall quality across the segments went 2.78 → 3.12 (+0.34) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This chain comes from the two-sided rule: it only counts if both emotions move — Bitterness down and Interest up — by at least 0.25 each.
The chain starts with Interest clearly present — 0.70, higher than 70 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.28.
At the same time Bitterness goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.68 (higher than 68 % of clips in this corpus), a change of -0.29. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.21, then +0.07 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.92 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.91 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.92), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.888 before conversion and 0.893 after — it rose by 0.005. Neighbour-to-neighbour the worst pair went 0.867 → 0.930. (The earlier render, with segment 1 left raw, scores 0.778 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Interest moved +0.278 in the original and +0.251 after conversion — 90 % of the delta retained, which is essentially all of it. On the other named axis, Bitterness, -0.292 became -0.644.
Quality. Mean predicted overall quality across the segments went 3.07 → 3.24 (+0.17) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This chain comes from the two-sided rule: it only counts if both emotions move — Jealousy and Envy down and Doubt up — by at least 0.25 each.
The chain starts with Doubt around average — 0.55, higher than 55 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.42.
At the same time Jealousy and Envy goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.72 (higher than 72 % of clips in this corpus), a change of -0.28. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.22, then +0.20 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.89 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.93 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.89. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.757 before conversion and 0.685 after — it fell by 0.071. Neighbour-to-neighbour the worst pair went 0.757 → 0.711. (The earlier render, with segment 1 left raw, scores 0.585 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Doubt moved +0.419 in the original and +0.455 after conversion — 109 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Jealousy and Envy, -0.276 became -0.352.
Quality. Mean predicted overall quality across the segments went 2.87 → 3.00 (+0.13) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This chain comes from the two-sided rule: it only counts if both emotions move — Concentration down and Emotional Numbness up — by at least 0.25 each.
The chain starts with Emotional Numbness around average — 0.50, right about the corpus median — and ends with it strongly present at 0.87, higher than 87 % of clips in this corpus. That is a total rise of 0.36.
At the same time Concentration goes the other way, from 0.84 (higher than 84 % of clips in this corpus) to 0.42 (lower than 58 % of clips in this corpus), a change of -0.42. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.21, then +0.15 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.98 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.98 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.98), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.922 before conversion and 0.855 after — it fell by 0.067. Neighbour-to-neighbour the worst pair went 0.933 → 0.888. (The earlier render, with segment 1 left raw, scores 0.608 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.361 in the original and +0.349 after conversion — 96 % of the delta retained, which is essentially all of it. On the other named axis, Concentration, -0.423 became -0.417.
Quality. Mean predicted overall quality across the segments went 2.91 → 3.05 (+0.14) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This chain comes from the two-sided rule: it only counts if both emotions move — Contemplation down and Fatigue Exhaustion up — by at least 0.25 each.
The chain starts with Fatigue Exhaustion clearly present — 0.59, higher than 59 % of clips in this corpus — and ends with it strongly present at 0.86, higher than 86 % of clips in this corpus. That is a total rise of 0.26.
At the same time Contemplation goes the other way, from 0.90 (higher than 90 % of clips in this corpus) to 0.57 (higher than 57 % of clips in this corpus), a change of -0.33. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.21, then +0.06 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.90 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.87 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.90), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.861 before conversion and 0.899 after — it rose by 0.038. Neighbour-to-neighbour the worst pair went 0.879 → 0.910. (The earlier render, with segment 1 left raw, scores 0.764 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fatigue Exhaustion moved +0.288 in the original and +0.347 after conversion — 121 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Contemplation, -0.330 became -0.459.
Quality. Mean predicted overall quality across the segments went 3.12 → 3.25 (+0.13) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This chain comes from the two-sided rule: it only counts if both emotions move — Pride down and Interest up — by at least 0.25 each.
The chain starts with Interest clearly present — 0.61, higher than 61 % of clips in this corpus — and ends with it strongly present at 0.87, higher than 87 % of clips in this corpus. That is a total rise of 0.26.
At the same time Pride goes the other way, from 0.94 (higher than 94 % of clips in this corpus) to 0.53 (higher than 53 % of clips in this corpus), a change of -0.41. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.10, then +0.16 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.94 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.95 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.94), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.942 before conversion and 0.931 after — it fell by 0.012. Neighbour-to-neighbour the worst pair went 0.956 → 0.931. (The earlier render, with segment 1 left raw, scores 0.889 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Interest moved +0.261 in the original and +0.068 after conversion — 26 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Pride, -0.416 became -0.381.
Quality. Mean predicted overall quality across the segments went 3.12 → 3.37 (+0.25) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This chain comes from the two-sided rule: it only counts if both emotions move — Disgust down and Astonishment Surprise up — by at least 0.25 each.
The chain starts with Astonishment Surprise clearly present — 0.66, higher than 66 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.32.
At the same time Disgust goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.65 (higher than 65 % of clips in this corpus), a change of -0.34. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.15, then +0.17 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.67 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.73 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.67, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.629 before conversion and 0.482 after — it fell by 0.147. Neighbour-to-neighbour the worst pair went 0.629 → 0.482. (The earlier render, with segment 1 left raw, scores 0.532 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Astonishment Surprise moved +0.455 in the original and +0.266 after conversion — 58 % of the delta retained. On the other named axis, Disgust, -0.334 became -0.218.
Quality. Mean predicted overall quality across the segments went 2.55 → 2.92 (+0.37) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This chain comes from the two-sided rule: it only counts if both emotions move — Impatience and Irritability down and Doubt up — by at least 0.25 each.
The chain starts with Doubt clearly present — 0.64, higher than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.31.
At the same time Impatience and Irritability goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.74 (higher than 74 % of clips in this corpus), a change of -0.25. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.16, then +0.15 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.88 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.87 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.88. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.841 before conversion and 0.799 after — it fell by 0.042. Neighbour-to-neighbour the worst pair went 0.841 → 0.799. (The earlier render, with segment 1 left raw, scores 0.748 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Doubt moved +0.311 in the original and +0.427 after conversion — 138 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Impatience and Irritability, -0.251 became -0.037.
Quality. Mean predicted overall quality across the segments went 2.80 → 2.96 (+0.16) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This chain comes from the two-sided rule: it only counts if both emotions move — Emotional Numbness down and Concentration up — by at least 0.25 each.
The chain starts with Concentration clearly present — 0.68, higher than 68 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.30.
At the same time Emotional Numbness goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.75 (higher than 75 % of clips in this corpus), a change of -0.25. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.13, then +0.16 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.94 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.94 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.94), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.913 before conversion and 0.903 after — it fell by 0.010. Neighbour-to-neighbour the worst pair went 0.893 → 0.872. (The earlier render, with segment 1 left raw, scores 0.817 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.297 in the original and +0.206 after conversion — 70 % of the delta retained. On the other named axis, Emotional Numbness, -0.252 became -0.296.
Quality. Mean predicted overall quality across the segments went 3.07 → 3.27 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This chain comes from the two-sided rule: it only counts if both emotions move — Teasing down and Doubt up — by at least 0.25 each.
The chain starts with Doubt clearly present — 0.64, higher than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.36.
At the same time Teasing goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.74 (higher than 74 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.15, then +0.20 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.17 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.55 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.17, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.208 before conversion and 0.736 after — it rose by 0.527. Neighbour-to-neighbour the worst pair went 0.481 → 0.759. (The earlier render, with segment 1 left raw, scores 0.425 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Doubt moved +0.358 in the original and +0.190 after conversion — 53 % of the delta retained. On the other named axis, Teasing, -0.258 became -0.635.
Quality. Mean predicted overall quality across the segments went 2.80 → 3.03 (+0.23) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This chain comes from the two-sided rule: it only counts if both emotions move — Hope Enthusiasm Optimism down and Interest up — by at least 0.25 each.
The chain starts with Interest around average — 0.52, higher than 52 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.47.
At the same time Hope Enthusiasm Optimism goes the other way, from 0.68 (higher than 68 % of clips in this corpus) to 0.97 (higher than 97 % of clips in this corpus), a change of +0.29. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.23, then +0.24 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.79 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.80 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.79, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.678 before conversion and 0.623 after — it fell by 0.055. Neighbour-to-neighbour the worst pair went 0.708 → 0.639. (The earlier render, with segment 1 left raw, scores 0.536 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Interest moved +0.470 in the original and +0.732 after conversion — 156 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Hope Enthusiasm Optimism, +0.289 became +0.350.
Quality. Mean predicted overall quality across the segments went 2.36 → 2.86 (+0.50) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This chain comes from the two-sided rule: it only counts if both emotions move — Sexual Lust down and Relief up — by at least 0.25 each.
The chain starts with Relief around average — 0.50, right about the corpus median — and ends with it strongly present at 0.87, higher than 87 % of clips in this corpus. That is a total rise of 0.37.
At the same time Sexual Lust goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.72 (higher than 72 % of clips in this corpus), a change of -0.27. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.22, then +0.15 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.73 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.81 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.73, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.709 before conversion and 0.724 after — it rose by 0.015. Neighbour-to-neighbour the worst pair went 0.663 → 0.626. (The earlier render, with segment 1 left raw, scores 0.632 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Relief moved +0.368 in the original and +0.044 after conversion — 12 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Sexual Lust, -0.270 became -0.051.
Quality. Mean predicted overall quality across the segments went 2.81 → 2.94 (+0.13) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This chain comes from the two-sided rule: it only counts if both emotions move — Infatuation down and Relief up — by at least 0.25 each.
The chain starts with Relief around average — 0.50, right about the corpus median — and ends with it strongly present at 0.77, higher than 77 % of clips in this corpus. That is a total rise of 0.27.
At the same time Infatuation goes the other way, from 0.89 (higher than 89 % of clips in this corpus) to 0.51 (right about the corpus median), a change of -0.38. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.24, then +0.03 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.94 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.94 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.94), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.915 before conversion and 0.891 after — it fell by 0.024. Neighbour-to-neighbour the worst pair went 0.915 → 0.891. (The earlier render, with segment 1 left raw, scores 0.908 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Relief moved +0.273 in the original and +0.119 after conversion — 43 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Infatuation, -0.378 became -0.374.
Quality. Mean predicted overall quality across the segments went 3.23 → 3.17 (-0.06) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This chain comes from the two-sided rule: it only counts if both emotions move — Contemplation down and Hope Enthusiasm Optimism up — by at least 0.25 each.
The chain starts with Hope Enthusiasm Optimism around average — 0.52, higher than 52 % of clips in this corpus — and ends with it strongly present at 0.87, higher than 87 % of clips in this corpus. That is a total rise of 0.36.
At the same time Contemplation goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.54 (higher than 54 % of clips in this corpus), a change of -0.44. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.16, then +0.20 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.73 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.69 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.73, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.749 before conversion and 0.616 after — it fell by 0.133. Neighbour-to-neighbour the worst pair went 0.717 → 0.616. (The earlier render, with segment 1 left raw, scores 0.496 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Hope Enthusiasm Optimism moved +0.357 in the original and +0.602 after conversion — 169 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Contemplation, -0.438 became -0.811.
Quality. Mean predicted overall quality across the segments went 2.25 → 2.84 (+0.59) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This chain comes from the two-sided rule: it only counts if both emotions move — Doubt down and Concentration up — by at least 0.25 each.
The chain starts with Concentration clearly present — 0.62, higher than 62 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.31.
At the same time Doubt goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.69 (higher than 69 % of clips in this corpus), a change of -0.30. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.11, then +0.21 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores -0.11 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.06 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (-0.11, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was -0.099 before conversion and 0.278 after — it rose by 0.377. Neighbour-to-neighbour the worst pair went 0.113 → 0.390. (The earlier render, with segment 1 left raw, scores 0.143 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.312 in the original and +0.453 after conversion — 145 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Doubt, -0.300 became -0.397.
Quality. Mean predicted overall quality across the segments went 2.69 → 3.07 (+0.37) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.