Sadness under rescue rule S4, k=2. one hop straight across the gap, step cap lifted. 90.0 % of clips score at or below zero on this emotion and the largest gap on its normalised axis is 0.449 (WIDER than the 0.25 step cap). This rule found 19,060 chains over 40,000 tracks; the strict rule found 0 at k=3.
S4 changes: Same total move, bigger allowed step. Keep the 0.25 end-to-end requirement on the normalised scale but lift the per-step cap so a chain may cross the gap in one hop. What it costs: The chain is no longer a smooth ramp. It is allowed to jump the whole distance in a single step, which is exactly what the step cap existed to forbid. Use it for k=2 pairs, where there is only one step anyway, and be careful reading it as a gradual progression. Full explanation →Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
This is not a strict-rule trajectory. It comes from rescue rule S4, which exists because the strict rule returns nothing at all for Sadness.
The raw scorer output across the chain is -0.00, 0.44. In the first clip the scorer found no Sadness whatsoever (-0.00); by the last it is at 0.44.
On the corpus-wide percentile scale those become 0.43, 0.92 — a total move of +0.49.
That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.
What this rule changes: Same total move, bigger allowed step. Keep the 0.25 end-to-end requirement on the normalised scale but lift the per-step cap so a chain may cross the gap in one hop.
What it costs: The chain is no longer a smooth ramp. It is allowed to jump the whole distance in a single step, which is exactly what the step cap existed to forbid. Use it for k=2 pairs, where there is only one step anyway, and be careful reading it as a gradual progression.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.747 before conversion and 0.776 after — it rose by 0.028. Neighbour-to-neighbour the worst pair went 0.747 → 0.776. (The earlier render, with segment 1 left raw, scores 0.691 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.486 in the original and +0.537 after conversion — 110 % of the delta retained, i.e. the move came out slightly larger after conversion than before.
Quality. Mean predicted overall quality across the segments went 2.94 → 3.19 (+0.24) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This is not a strict-rule trajectory. It comes from rescue rule S4, which exists because the strict rule returns nothing at all for Sadness.
The raw scorer output across the chain is -0.00, 0.79. In the first clip the scorer found no Sadness whatsoever (-0.00); by the last it is at 0.79.
On the corpus-wide percentile scale those become 0.43, 0.96 — a total move of +0.53.
That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.
What this rule changes: Same total move, bigger allowed step. Keep the 0.25 end-to-end requirement on the normalised scale but lift the per-step cap so a chain may cross the gap in one hop.
What it costs: The chain is no longer a smooth ramp. It is allowed to jump the whole distance in a single step, which is exactly what the step cap existed to forbid. Use it for k=2 pairs, where there is only one step anyway, and be careful reading it as a gradual progression.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.308 before conversion and 0.797 after — it rose by 0.489. Neighbour-to-neighbour the worst pair went 0.308 → 0.797. (The earlier render, with segment 1 left raw, scores 0.740 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.528 in the original and +0.508 after conversion — 96 % of the delta retained, which is essentially all of it.
Quality. Mean predicted overall quality across the segments went 2.82 → 3.10 (+0.28) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This is not a strict-rule trajectory. It comes from rescue rule S4, which exists because the strict rule returns nothing at all for Sadness.
The raw scorer output across the chain is -0.00, 0.41. In the first clip the scorer found no Sadness whatsoever (-0.00); by the last it is at 0.41.
On the corpus-wide percentile scale those become 0.43, 0.91 — a total move of +0.48.
That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.
What this rule changes: Same total move, bigger allowed step. Keep the 0.25 end-to-end requirement on the normalised scale but lift the per-step cap so a chain may cross the gap in one hop.
What it costs: The chain is no longer a smooth ramp. It is allowed to jump the whole distance in a single step, which is exactly what the step cap existed to forbid. Use it for k=2 pairs, where there is only one step anyway, and be careful reading it as a gradual progression.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.777 before conversion and 0.796 after — it rose by 0.019. Neighbour-to-neighbour the worst pair went 0.777 → 0.796. (The earlier render, with segment 1 left raw, scores 0.756 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.483 in the original and +0.520 after conversion — 108 % of the delta retained, i.e. the move came out slightly larger after conversion than before.
Quality. Mean predicted overall quality across the segments went 2.96 → 3.09 (+0.13) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This is not a strict-rule trajectory. It comes from rescue rule S4, which exists because the strict rule returns nothing at all for Sadness.
The raw scorer output across the chain is -0.00, 0.74. In the first clip the scorer found no Sadness whatsoever (-0.00); by the last it is at 0.74.
On the corpus-wide percentile scale those become 0.43, 0.95 — a total move of +0.52.
That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.
What this rule changes: Same total move, bigger allowed step. Keep the 0.25 end-to-end requirement on the normalised scale but lift the per-step cap so a chain may cross the gap in one hop.
What it costs: The chain is no longer a smooth ramp. It is allowed to jump the whole distance in a single step, which is exactly what the step cap existed to forbid. Use it for k=2 pairs, where there is only one step anyway, and be careful reading it as a gradual progression.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was -0.023 before conversion and 0.618 after — it rose by 0.641. Neighbour-to-neighbour the worst pair went -0.023 → 0.618. (The earlier render, with segment 1 left raw, scores 0.529 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.523 in the original and +0.027 after conversion — 5 % of the delta retained, so a meaningful part of the trajectory was flattened.
Quality. Mean predicted overall quality across the segments went 2.93 → 3.17 (+0.25) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This is not a strict-rule trajectory. It comes from rescue rule S4, which exists because the strict rule returns nothing at all for Sadness.
The raw scorer output across the chain is 0.00, 0.49. In the first clip the scorer found no Sadness whatsoever (0.00); by the last it is at 0.49.
On the corpus-wide percentile scale those become 0.43, 0.92 — a total move of +0.49.
That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.
What this rule changes: Same total move, bigger allowed step. Keep the 0.25 end-to-end requirement on the normalised scale but lift the per-step cap so a chain may cross the gap in one hop.
What it costs: The chain is no longer a smooth ramp. It is allowed to jump the whole distance in a single step, which is exactly what the step cap existed to forbid. Use it for k=2 pairs, where there is only one step anyway, and be careful reading it as a gradual progression.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.408 before conversion and 0.678 after — it rose by 0.270. Neighbour-to-neighbour the worst pair went 0.408 → 0.678. (The earlier render, with segment 1 left raw, scores 0.507 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.492 in the original and +0.000 after conversion — 0 % of the delta retained, so a meaningful part of the trajectory was flattened.
Quality. Mean predicted overall quality across the segments went 3.01 → 3.16 (+0.14) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This is not a strict-rule trajectory. It comes from rescue rule S4, which exists because the strict rule returns nothing at all for Sadness.
The raw scorer output across the chain is 0.00, 0.00. In the first clip the scorer found no Sadness whatsoever (0.00); by the last it is at 0.00.
On the corpus-wide percentile scale those become 0.43, 0.86 — a total move of +0.43.
That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.
What this rule changes: Same total move, bigger allowed step. Keep the 0.25 end-to-end requirement on the normalised scale but lift the per-step cap so a chain may cross the gap in one hop.
What it costs: The chain is no longer a smooth ramp. It is allowed to jump the whole distance in a single step, which is exactly what the step cap existed to forbid. Use it for k=2 pairs, where there is only one step anyway, and be careful reading it as a gradual progression.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.751 before conversion and 0.762 after — it rose by 0.011. Neighbour-to-neighbour the worst pair went 0.751 → 0.762. (The earlier render, with segment 1 left raw, scores 0.677 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.430 in the original and +0.000 after conversion — 0 % of the delta retained, so a meaningful part of the trajectory was flattened.
Quality. Mean predicted overall quality across the segments went 2.70 → 2.90 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This is not a strict-rule trajectory. It comes from rescue rule S4, which exists because the strict rule returns nothing at all for Sadness.
The raw scorer output across the chain is -0.00, 0.38. In the first clip the scorer found no Sadness whatsoever (-0.00); by the last it is at 0.38.
On the corpus-wide percentile scale those become 0.43, 0.91 — a total move of +0.48.
That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.
What this rule changes: Same total move, bigger allowed step. Keep the 0.25 end-to-end requirement on the normalised scale but lift the per-step cap so a chain may cross the gap in one hop.
What it costs: The chain is no longer a smooth ramp. It is allowed to jump the whole distance in a single step, which is exactly what the step cap existed to forbid. Use it for k=2 pairs, where there is only one step anyway, and be careful reading it as a gradual progression.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.514 before conversion and 0.736 after — it rose by 0.221. Neighbour-to-neighbour the worst pair went 0.514 → 0.736. (The earlier render, with segment 1 left raw, scores 0.654 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.479 in the original and +0.506 after conversion — 106 % of the delta retained, i.e. the move came out slightly larger after conversion than before.
Quality. Mean predicted overall quality across the segments went 2.94 → 3.11 (+0.17) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This is not a strict-rule trajectory. It comes from rescue rule S4, which exists because the strict rule returns nothing at all for Sadness.
The raw scorer output across the chain is -0.00, 0.66. In the first clip the scorer found no Sadness whatsoever (-0.00); by the last it is at 0.66.
On the corpus-wide percentile scale those become 0.43, 0.94 — a total move of +0.51.
That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.
What this rule changes: Same total move, bigger allowed step. Keep the 0.25 end-to-end requirement on the normalised scale but lift the per-step cap so a chain may cross the gap in one hop.
What it costs: The chain is no longer a smooth ramp. It is allowed to jump the whole distance in a single step, which is exactly what the step cap existed to forbid. Use it for k=2 pairs, where there is only one step anyway, and be careful reading it as a gradual progression.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.900 before conversion and 0.924 after — it rose by 0.023. Neighbour-to-neighbour the worst pair went 0.900 → 0.924. (The earlier render, with segment 1 left raw, scores 0.869 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.514 in the original and +0.441 after conversion — 86 % of the delta retained, which is most of it.
Quality. Mean predicted overall quality across the segments went 2.86 → 3.15 (+0.29) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This is not a strict-rule trajectory. It comes from rescue rule S4, which exists because the strict rule returns nothing at all for Sadness.
The raw scorer output across the chain is -0.00, 0.23. In the first clip the scorer found no Sadness whatsoever (-0.00); by the last it is at 0.23.
On the corpus-wide percentile scale those become 0.43, 0.89 — a total move of +0.46.
That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.
What this rule changes: Same total move, bigger allowed step. Keep the 0.25 end-to-end requirement on the normalised scale but lift the per-step cap so a chain may cross the gap in one hop.
What it costs: The chain is no longer a smooth ramp. It is allowed to jump the whole distance in a single step, which is exactly what the step cap existed to forbid. Use it for k=2 pairs, where there is only one step anyway, and be careful reading it as a gradual progression.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.850 before conversion and 0.872 after — it rose by 0.022. Neighbour-to-neighbour the worst pair went 0.850 → 0.872. (The earlier render, with segment 1 left raw, scores 0.815 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.463 in the original and +0.550 after conversion — 119 % of the delta retained, i.e. the move came out slightly larger after conversion than before.
Quality. Mean predicted overall quality across the segments went 2.91 → 3.08 (+0.18) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This is not a strict-rule trajectory. It comes from rescue rule S4, which exists because the strict rule returns nothing at all for Sadness.
The raw scorer output across the chain is -0.00, 0.09. In the first clip the scorer found no Sadness whatsoever (-0.00); by the last it is at 0.09.
On the corpus-wide percentile scale those become 0.43, 0.88 — a total move of +0.45.
That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.
What this rule changes: Same total move, bigger allowed step. Keep the 0.25 end-to-end requirement on the normalised scale but lift the per-step cap so a chain may cross the gap in one hop.
What it costs: The chain is no longer a smooth ramp. It is allowed to jump the whole distance in a single step, which is exactly what the step cap existed to forbid. Use it for k=2 pairs, where there is only one step anyway, and be careful reading it as a gradual progression.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.878 before conversion and 0.908 after — it rose by 0.029. Neighbour-to-neighbour the worst pair went 0.878 → 0.908. (The earlier render, with segment 1 left raw, scores 0.870 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.449 in the original and +0.000 after conversion — 0 % of the delta retained, so a meaningful part of the trajectory was flattened.
Quality. Mean predicted overall quality across the segments went 2.85 → 3.20 (+0.35) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This is not a strict-rule trajectory. It comes from rescue rule S4, which exists because the strict rule returns nothing at all for Sadness.
The raw scorer output across the chain is 0.00, 0.05. In the first clip the scorer found no Sadness whatsoever (0.00); by the last it is at 0.05.
On the corpus-wide percentile scale those become 0.43, 0.87 — a total move of +0.44.
That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.
What this rule changes: Same total move, bigger allowed step. Keep the 0.25 end-to-end requirement on the normalised scale but lift the per-step cap so a chain may cross the gap in one hop.
What it costs: The chain is no longer a smooth ramp. It is allowed to jump the whole distance in a single step, which is exactly what the step cap existed to forbid. Use it for k=2 pairs, where there is only one step anyway, and be careful reading it as a gradual progression.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.772 before conversion and 0.821 after — it rose by 0.049. Neighbour-to-neighbour the worst pair went 0.772 → 0.821. (The earlier render, with segment 1 left raw, scores 0.731 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.443 in the original and +0.000 after conversion — 0 % of the delta retained, so a meaningful part of the trajectory was flattened.
Quality. Mean predicted overall quality across the segments went 2.71 → 3.06 (+0.35) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This is not a strict-rule trajectory. It comes from rescue rule S4, which exists because the strict rule returns nothing at all for Sadness.
The raw scorer output across the chain is 0.00, 0.26. In the first clip the scorer found no Sadness whatsoever (0.00); by the last it is at 0.26.
On the corpus-wide percentile scale those become 0.43, 0.90 — a total move of +0.47.
That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.
What this rule changes: Same total move, bigger allowed step. Keep the 0.25 end-to-end requirement on the normalised scale but lift the per-step cap so a chain may cross the gap in one hop.
What it costs: The chain is no longer a smooth ramp. It is allowed to jump the whole distance in a single step, which is exactly what the step cap existed to forbid. Use it for k=2 pairs, where there is only one step anyway, and be careful reading it as a gradual progression.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.876 before conversion and 0.915 after — it rose by 0.039. Neighbour-to-neighbour the worst pair went 0.876 → 0.915. (The earlier render, with segment 1 left raw, scores 0.881 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.466 in the original and +0.496 after conversion — 106 % of the delta retained, i.e. the move came out slightly larger after conversion than before.
Quality. Mean predicted overall quality across the segments went 2.74 → 3.18 (+0.44) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This is not a strict-rule trajectory. It comes from rescue rule S4, which exists because the strict rule returns nothing at all for Sadness.
The raw scorer output across the chain is -0.00, 0.07. In the first clip the scorer found no Sadness whatsoever (-0.00); by the last it is at 0.07.
On the corpus-wide percentile scale those become 0.43, 0.88 — a total move of +0.45.
That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.
What this rule changes: Same total move, bigger allowed step. Keep the 0.25 end-to-end requirement on the normalised scale but lift the per-step cap so a chain may cross the gap in one hop.
What it costs: The chain is no longer a smooth ramp. It is allowed to jump the whole distance in a single step, which is exactly what the step cap existed to forbid. Use it for k=2 pairs, where there is only one step anyway, and be careful reading it as a gradual progression.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.832 before conversion and 0.843 after — it rose by 0.010. Neighbour-to-neighbour the worst pair went 0.832 → 0.843. (The earlier render, with segment 1 left raw, scores 0.793 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.447 in the original and +0.000 after conversion — 0 % of the delta retained, so a meaningful part of the trajectory was flattened.
Quality. Mean predicted overall quality across the segments went 3.19 → 3.37 (+0.18) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This is not a strict-rule trajectory. It comes from rescue rule S4, which exists because the strict rule returns nothing at all for Sadness.
The raw scorer output across the chain is 0.00, 1.33. In the first clip the scorer found no Sadness whatsoever (0.00); by the last it is at 1.33.
On the corpus-wide percentile scale those become 0.39, 0.99 — a total move of +0.59.
That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.
What this rule changes: Same total move, bigger allowed step. Keep the 0.25 end-to-end requirement on the normalised scale but lift the per-step cap so a chain may cross the gap in one hop.
What it costs: The chain is no longer a smooth ramp. It is allowed to jump the whole distance in a single step, which is exactly what the step cap existed to forbid. Use it for k=2 pairs, where there is only one step anyway, and be careful reading it as a gradual progression.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.897 before conversion and 0.909 after — it rose by 0.012. Neighbour-to-neighbour the worst pair went 0.897 → 0.909. (The earlier render, with segment 1 left raw, scores 0.792 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.474 in the original and +0.434 after conversion — 92 % of the delta retained, which is essentially all of it.
Quality. Mean predicted overall quality across the segments went 2.82 → 3.16 (+0.34) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This is not a strict-rule trajectory. It comes from rescue rule S4, which exists because the strict rule returns nothing at all for Sadness.
The raw scorer output across the chain is -0.00, 0.70. In the first clip the scorer found no Sadness whatsoever (-0.00); by the last it is at 0.70.
On the corpus-wide percentile scale those become 0.43, 0.95 — a total move of +0.52.
That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.
What this rule changes: Same total move, bigger allowed step. Keep the 0.25 end-to-end requirement on the normalised scale but lift the per-step cap so a chain may cross the gap in one hop.
What it costs: The chain is no longer a smooth ramp. It is allowed to jump the whole distance in a single step, which is exactly what the step cap existed to forbid. Use it for k=2 pairs, where there is only one step anyway, and be careful reading it as a gradual progression.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.779 before conversion and 0.851 after — it rose by 0.072. Neighbour-to-neighbour the worst pair went 0.779 → 0.851. (The earlier render, with segment 1 left raw, scores 0.739 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.518 in the original and +0.535 after conversion — 103 % of the delta retained, which is essentially all of it.
Quality. Mean predicted overall quality across the segments went 2.65 → 2.93 (+0.28) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This is not a strict-rule trajectory. It comes from rescue rule S4, which exists because the strict rule returns nothing at all for Sadness.
The raw scorer output across the chain is -0.00, 0.36. In the first clip the scorer found no Sadness whatsoever (-0.00); by the last it is at 0.36.
On the corpus-wide percentile scale those become 0.43, 0.91 — a total move of +0.48.
That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.
What this rule changes: Same total move, bigger allowed step. Keep the 0.25 end-to-end requirement on the normalised scale but lift the per-step cap so a chain may cross the gap in one hop.
What it costs: The chain is no longer a smooth ramp. It is allowed to jump the whole distance in a single step, which is exactly what the step cap existed to forbid. Use it for k=2 pairs, where there is only one step anyway, and be careful reading it as a gradual progression.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.097 before conversion and 0.885 after — it rose by 0.788. Neighbour-to-neighbour the worst pair went 0.097 → 0.885. (The earlier render, with segment 1 left raw, scores 0.781 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.476 in the original and +0.443 after conversion — 93 % of the delta retained, which is essentially all of it.
Quality. Mean predicted overall quality across the segments went 3.04 → 3.25 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This is not a strict-rule trajectory. It comes from rescue rule S4, which exists because the strict rule returns nothing at all for Sadness.
The raw scorer output across the chain is -0.00, 0.00. In the first clip the scorer found no Sadness whatsoever (-0.00); by the last it is at 0.00.
On the corpus-wide percentile scale those become 0.43, 0.86 — a total move of +0.43.
That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.
What this rule changes: Same total move, bigger allowed step. Keep the 0.25 end-to-end requirement on the normalised scale but lift the per-step cap so a chain may cross the gap in one hop.
What it costs: The chain is no longer a smooth ramp. It is allowed to jump the whole distance in a single step, which is exactly what the step cap existed to forbid. Use it for k=2 pairs, where there is only one step anyway, and be careful reading it as a gradual progression.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.828 before conversion and 0.827 after — it fell by 0.001. Neighbour-to-neighbour the worst pair went 0.828 → 0.827. (The earlier render, with segment 1 left raw, scores 0.806 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.428 in the original and +0.000 after conversion — 0 % of the delta retained, so a meaningful part of the trajectory was flattened.
Quality. Mean predicted overall quality across the segments went 2.89 → 3.09 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This is not a strict-rule trajectory. It comes from rescue rule S4, which exists because the strict rule returns nothing at all for Sadness.
The raw scorer output across the chain is 0.00, 0.73. In the first clip the scorer found no Sadness whatsoever (0.00); by the last it is at 0.73.
On the corpus-wide percentile scale those become 0.43, 0.95 — a total move of +0.52.
That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.
What this rule changes: Same total move, bigger allowed step. Keep the 0.25 end-to-end requirement on the normalised scale but lift the per-step cap so a chain may cross the gap in one hop.
What it costs: The chain is no longer a smooth ramp. It is allowed to jump the whole distance in a single step, which is exactly what the step cap existed to forbid. Use it for k=2 pairs, where there is only one step anyway, and be careful reading it as a gradual progression.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.242 before conversion and 0.560 after — it rose by 0.318. Neighbour-to-neighbour the worst pair went 0.242 → 0.560. (The earlier render, with segment 1 left raw, scores 0.500 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.522 in the original and +0.518 after conversion — 99 % of the delta retained, which is essentially all of it.
Quality. Mean predicted overall quality across the segments went 2.66 → 3.11 (+0.44) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This is not a strict-rule trajectory. It comes from rescue rule S4, which exists because the strict rule returns nothing at all for Sadness.
The raw scorer output across the chain is -0.00, 0.00. In the first clip the scorer found no Sadness whatsoever (-0.00); by the last it is at 0.00.
On the corpus-wide percentile scale those become 0.43, 0.86 — a total move of +0.43.
That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.
What this rule changes: Same total move, bigger allowed step. Keep the 0.25 end-to-end requirement on the normalised scale but lift the per-step cap so a chain may cross the gap in one hop.
What it costs: The chain is no longer a smooth ramp. It is allowed to jump the whole distance in a single step, which is exactly what the step cap existed to forbid. Use it for k=2 pairs, where there is only one step anyway, and be careful reading it as a gradual progression.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.798 before conversion and 0.766 after — it fell by 0.031. Neighbour-to-neighbour the worst pair went 0.798 → 0.766. (The earlier render, with segment 1 left raw, scores 0.761 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.430 in the original and +0.000 after conversion — 0 % of the delta retained, so a meaningful part of the trajectory was flattened.
Quality. Mean predicted overall quality across the segments went 2.73 → 3.07 (+0.34) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This is not a strict-rule trajectory. It comes from rescue rule S4, which exists because the strict rule returns nothing at all for Sadness.
The raw scorer output across the chain is 0.00, 0.17. In the first clip the scorer found no Sadness whatsoever (0.00); by the last it is at 0.17.
On the corpus-wide percentile scale those become 0.43, 0.89 — a total move of +0.46.
That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Sadness, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Sadness sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.
What this rule changes: Same total move, bigger allowed step. Keep the 0.25 end-to-end requirement on the normalised scale but lift the per-step cap so a chain may cross the gap in one hop.
What it costs: The chain is no longer a smooth ramp. It is allowed to jump the whole distance in a single step, which is exactly what the step cap existed to forbid. Use it for k=2 pairs, where there is only one step anyway, and be careful reading it as a gradual progression.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.823 before conversion and 0.827 after — it rose by 0.005. Neighbour-to-neighbour the worst pair went 0.823 → 0.827. (The earlier render, with segment 1 left raw, scores 0.758 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.457 in the original and +0.482 after conversion — 106 % of the delta retained, i.e. the move came out slightly larger after conversion than before.
Quality. Mean predicted overall quality across the segments went 2.97 → 3.11 (+0.14) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.