Sadness under rescue rule S1, k=3. a genuine three-step ramp, once the zero block is removed. 90.0 % of clips score at or below zero on this emotion and the largest gap on its normalised axis is 0.449 (WIDER than the 0.25 step cap). This rule found 3,162 chains over 40,000 tracks; the strict rule found 0 at k=3.
S1 changes: Positives-only renormalisation. Throw away the huge block of clips that score zero -- 'no sadness' is not a degree of sadness -- and rank-normalise only the clips that actually carry the emotion. The gap disappears, and the ordinary 0.25 / 0.25 test then works normally. What it costs: The number stops being comparable with other emotions: 'moved 0.25' now means a quarter of the sadness-BEARING subpopulation, not a quarter of the corpus. The chains also come from a small, self-selected slice of the data. Full explanation →Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
This is not a strict-rule trajectory. It comes from rescue rule S1, which exists because the strict rule returns nothing at all for Sadness.
The raw scorer output across the chain is 0.33, 0.53, 0.70. In the first clip the scorer found only a trace of Sadness (0.33); by the last it is at 0.70.
On the corpus-wide percentile scale those become 0.90, 0.93, 0.95 — a total move of +0.04.
That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Sadness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.04 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.
Re-ranked among the Sadness-bearing clips only — which is exactly what this rule does — the same clips read 0.23, 0.40, 0.56, a move of +0.32, which is usable again.
What this rule changes: Positives-only renormalisation. Throw away the huge block of clips that score zero -- 'no sadness' is not a degree of sadness -- and rank-normalise only the clips that actually carry the emotion. The gap disappears, and the ordinary 0.25 / 0.25 test then works normally.
What it costs: The number stops being comparable with other emotions: 'moved 0.25' now means a quarter of the sadness-BEARING subpopulation, not a quarter of the corpus. The chains also come from a small, self-selected slice of the data.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.772 before conversion and 0.775 after — it rose by 0.003. Neighbour-to-neighbour the worst pair went 0.772 → 0.834. (The earlier render, with segment 1 left raw, scores 0.738 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.044 in the original and +0.535 after conversion — 1209 % of the delta retained, i.e. the move came out slightly larger after conversion than before.
Quality. Mean predicted overall quality across the segments went 2.83 → 3.08 (+0.25) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This is not a strict-rule trajectory. It comes from rescue rule S1, which exists because the strict rule returns nothing at all for Sadness.
The raw scorer output across the chain is 0.13, 0.40, 0.60. In the first clip the scorer found only a trace of Sadness (0.13); by the last it is at 0.60.
On the corpus-wide percentile scale those become 0.88, 0.91, 0.94 — a total move of +0.05.
That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Sadness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.05 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.
Re-ranked among the Sadness-bearing clips only — which is exactly what this rule does — the same clips read 0.08, 0.29, 0.47, a move of +0.39, which is usable again.
What this rule changes: Positives-only renormalisation. Throw away the huge block of clips that score zero -- 'no sadness' is not a degree of sadness -- and rank-normalise only the clips that actually carry the emotion. The gap disappears, and the ordinary 0.25 / 0.25 test then works normally.
What it costs: The number stops being comparable with other emotions: 'moved 0.25' now means a quarter of the sadness-BEARING subpopulation, not a quarter of the corpus. The chains also come from a small, self-selected slice of the data.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.930 before conversion and 0.933 after — it rose by 0.003. Neighbour-to-neighbour the worst pair went 0.949 → 0.933. (The earlier render, with segment 1 left raw, scores 0.823 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.054 in the original and +0.028 after conversion — 53 % of the delta retained.
Quality. Mean predicted overall quality across the segments went 3.02 → 3.18 (+0.16) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This is not a strict-rule trajectory. It comes from rescue rule S1, which exists because the strict rule returns nothing at all for Sadness.
The raw scorer output across the chain is 0.39, 0.56, 0.73. In the first clip the scorer found only a trace of Sadness (0.39); by the last it is at 0.73.
On the corpus-wide percentile scale those become 0.91, 0.93, 0.95 — a total move of +0.04.
That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Sadness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.04 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.
Re-ranked among the Sadness-bearing clips only — which is exactly what this rule does — the same clips read 0.28, 0.43, 0.59, a move of +0.31, which is usable again.
What this rule changes: Positives-only renormalisation. Throw away the huge block of clips that score zero -- 'no sadness' is not a degree of sadness -- and rank-normalise only the clips that actually carry the emotion. The gap disappears, and the ordinary 0.25 / 0.25 test then works normally.
What it costs: The number stops being comparable with other emotions: 'moved 0.25' now means a quarter of the sadness-BEARING subpopulation, not a quarter of the corpus. The chains also come from a small, self-selected slice of the data.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.143 before conversion and 0.685 after — it rose by 0.541. Neighbour-to-neighbour the worst pair went 0.234 → 0.685. (The earlier render, with segment 1 left raw, scores 0.519 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.042 in the original and +0.010 after conversion — 24 % of the delta retained, so a meaningful part of the trajectory was flattened.
Quality. Mean predicted overall quality across the segments went 2.77 → 3.26 (+0.48) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This is not a strict-rule trajectory. It comes from rescue rule S1, which exists because the strict rule returns nothing at all for Sadness.
The raw scorer output across the chain is 0.07, 0.13, 0.40. In the first clip the scorer found only a trace of Sadness (0.07); by the last it is at 0.40.
On the corpus-wide percentile scale those become 0.88, 0.88, 0.91 — a total move of +0.04.
That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Sadness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.04 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.
Re-ranked among the Sadness-bearing clips only — which is exactly what this rule does — the same clips read 0.02, 0.08, 0.28, a move of +0.27, which is usable again.
What this rule changes: Positives-only renormalisation. Throw away the huge block of clips that score zero -- 'no sadness' is not a degree of sadness -- and rank-normalise only the clips that actually carry the emotion. The gap disappears, and the ordinary 0.25 / 0.25 test then works normally.
What it costs: The number stops being comparable with other emotions: 'moved 0.25' now means a quarter of the sadness-BEARING subpopulation, not a quarter of the corpus. The chains also come from a small, self-selected slice of the data.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.703 before conversion and 0.640 after — it fell by 0.063. Neighbour-to-neighbour the worst pair went 0.703 → 0.640. (The earlier render, with segment 1 left raw, scores 0.594 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.036 in the original and +0.000 after conversion — 0 % of the delta retained, so a meaningful part of the trajectory was flattened.
Quality. Mean predicted overall quality across the segments went 2.77 → 3.00 (+0.24) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This is not a strict-rule trajectory. It comes from rescue rule S1, which exists because the strict rule returns nothing at all for Sadness.
The raw scorer output across the chain is 0.83, 1.11, 1.39. In the first clip the scorer already found some Sadness here (0.83); by the last it is at 1.39.
On the corpus-wide percentile scale those become 0.96, 0.98, 0.99 — a total move of +0.03.
That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Sadness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.03 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.
Re-ranked among the Sadness-bearing clips only — which is exactly what this rule does — the same clips read 0.67, 0.84, 0.93, a move of +0.26, which is usable again.
What this rule changes: Positives-only renormalisation. Throw away the huge block of clips that score zero -- 'no sadness' is not a degree of sadness -- and rank-normalise only the clips that actually carry the emotion. The gap disappears, and the ordinary 0.25 / 0.25 test then works normally.
What it costs: The number stops being comparable with other emotions: 'moved 0.25' now means a quarter of the sadness-BEARING subpopulation, not a quarter of the corpus. The chains also come from a small, self-selected slice of the data.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.303 before conversion and 0.774 after — it rose by 0.470. Neighbour-to-neighbour the worst pair went 0.434 → 0.794. (The earlier render, with segment 1 left raw, scores 0.675 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.031 in the original and +0.075 after conversion — 245 % of the delta retained, i.e. the move came out slightly larger after conversion than before.
Quality. Mean predicted overall quality across the segments went 2.93 → 3.18 (+0.25) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This is not a strict-rule trajectory. It comes from rescue rule S1, which exists because the strict rule returns nothing at all for Sadness.
The raw scorer output across the chain is 0.47, 0.60, 0.79. In the first clip the scorer found only a trace of Sadness (0.47); by the last it is at 0.79.
On the corpus-wide percentile scale those become 0.92, 0.94, 0.96 — a total move of +0.04.
That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Sadness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.04 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.
Re-ranked among the Sadness-bearing clips only — which is exactly what this rule does — the same clips read 0.35, 0.46, 0.64, a move of +0.30, which is usable again.
What this rule changes: Positives-only renormalisation. Throw away the huge block of clips that score zero -- 'no sadness' is not a degree of sadness -- and rank-normalise only the clips that actually carry the emotion. The gap disappears, and the ordinary 0.25 / 0.25 test then works normally.
What it costs: The number stops being comparable with other emotions: 'moved 0.25' now means a quarter of the sadness-BEARING subpopulation, not a quarter of the corpus. The chains also come from a small, self-selected slice of the data.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.869 before conversion and 0.876 after — it rose by 0.007. Neighbour-to-neighbour the worst pair went 0.869 → 0.902. (The earlier render, with segment 1 left raw, scores 0.764 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
The emotional move did not survive. Re-scored end to end, Sadness moved +0.038 in the original and -0.002 after conversion — it changed direction. On this chain the corrected audio is not an improvement.
Quality. Mean predicted overall quality across the segments went 2.60 → 3.20 (+0.60) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This is not a strict-rule trajectory. It comes from rescue rule S1, which exists because the strict rule returns nothing at all for Sadness.
The raw scorer output across the chain is 0.24, 0.39, 0.64. In the first clip the scorer found only a trace of Sadness (0.24); by the last it is at 0.64.
On the corpus-wide percentile scale those become 0.89, 0.91, 0.94 — a total move of +0.05.
That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Sadness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.05 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.
Re-ranked among the Sadness-bearing clips only — which is exactly what this rule does — the same clips read 0.17, 0.28, 0.50, a move of +0.33, which is usable again.
What this rule changes: Positives-only renormalisation. Throw away the huge block of clips that score zero -- 'no sadness' is not a degree of sadness -- and rank-normalise only the clips that actually carry the emotion. The gap disappears, and the ordinary 0.25 / 0.25 test then works normally.
What it costs: The number stops being comparable with other emotions: 'moved 0.25' now means a quarter of the sadness-BEARING subpopulation, not a quarter of the corpus. The chains also come from a small, self-selected slice of the data.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.847 before conversion and 0.872 after — it rose by 0.025. Neighbour-to-neighbour the worst pair went 0.847 → 0.872. (The earlier render, with segment 1 left raw, scores 0.798 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
The emotional move did not survive. Re-scored end to end, Sadness moved +0.046 in the original and -0.021 after conversion — it changed direction. On this chain the corrected audio is not an improvement.
Quality. Mean predicted overall quality across the segments went 2.89 → 3.15 (+0.26) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This is not a strict-rule trajectory. It comes from rescue rule S1, which exists because the strict rule returns nothing at all for Sadness.
The raw scorer output across the chain is 0.17, 0.34, 0.61. In the first clip the scorer found only a trace of Sadness (0.17); by the last it is at 0.61.
On the corpus-wide percentile scale those become 0.89, 0.91, 0.94 — a total move of +0.05.
That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Sadness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.05 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.
Re-ranked among the Sadness-bearing clips only — which is exactly what this rule does — the same clips read 0.11, 0.24, 0.47, a move of +0.36, which is usable again.
What this rule changes: Positives-only renormalisation. Throw away the huge block of clips that score zero -- 'no sadness' is not a degree of sadness -- and rank-normalise only the clips that actually carry the emotion. The gap disappears, and the ordinary 0.25 / 0.25 test then works normally.
What it costs: The number stops being comparable with other emotions: 'moved 0.25' now means a quarter of the sadness-BEARING subpopulation, not a quarter of the corpus. The chains also come from a small, self-selected slice of the data.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.850 before conversion and 0.826 after — it fell by 0.025. Neighbour-to-neighbour the worst pair went 0.868 → 0.846. (The earlier render, with segment 1 left raw, scores 0.744 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.049 in the original and +0.499 after conversion — 1018 % of the delta retained, i.e. the move came out slightly larger after conversion than before.
Quality. Mean predicted overall quality across the segments went 3.07 → 3.25 (+0.18) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This is not a strict-rule trajectory. It comes from rescue rule S1, which exists because the strict rule returns nothing at all for Sadness.
The raw scorer output across the chain is 0.79, 1.08, 1.40. In the first clip the scorer already found some Sadness here (0.79); by the last it is at 1.40.
On the corpus-wide percentile scale those become 0.96, 0.98, 0.99 — a total move of +0.04.
That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Sadness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.04 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.
Re-ranked among the Sadness-bearing clips only — which is exactly what this rule does — the same clips read 0.64, 0.83, 0.93, a move of +0.30, which is usable again.
What this rule changes: Positives-only renormalisation. Throw away the huge block of clips that score zero -- 'no sadness' is not a degree of sadness -- and rank-normalise only the clips that actually carry the emotion. The gap disappears, and the ordinary 0.25 / 0.25 test then works normally.
What it costs: The number stops being comparable with other emotions: 'moved 0.25' now means a quarter of the sadness-BEARING subpopulation, not a quarter of the corpus. The chains also come from a small, self-selected slice of the data.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.581 before conversion and 0.765 after — it rose by 0.184. Neighbour-to-neighbour the worst pair went 0.759 → 0.802. (The earlier render, with segment 1 left raw, scores 0.641 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.035 in the original and +0.004 after conversion — 11 % of the delta retained, so a meaningful part of the trajectory was flattened.
Quality. Mean predicted overall quality across the segments went 2.55 → 2.94 (+0.40) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This is not a strict-rule trajectory. It comes from rescue rule S1, which exists because the strict rule returns nothing at all for Sadness.
The raw scorer output across the chain is 0.51, 0.66, 0.81. In the first clip the scorer already found some Sadness here (0.51); by the last it is at 0.81.
On the corpus-wide percentile scale those become 0.92, 0.94, 0.96 — a total move of +0.04.
That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Sadness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.04 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.
Re-ranked among the Sadness-bearing clips only — which is exactly what this rule does — the same clips read 0.38, 0.52, 0.65, a move of +0.28, which is usable again.
What this rule changes: Positives-only renormalisation. Throw away the huge block of clips that score zero -- 'no sadness' is not a degree of sadness -- and rank-normalise only the clips that actually carry the emotion. The gap disappears, and the ordinary 0.25 / 0.25 test then works normally.
What it costs: The number stops being comparable with other emotions: 'moved 0.25' now means a quarter of the sadness-BEARING subpopulation, not a quarter of the corpus. The chains also come from a small, self-selected slice of the data.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.855 before conversion and 0.816 after — it fell by 0.039. Neighbour-to-neighbour the worst pair went 0.821 → 0.742. (The earlier render, with segment 1 left raw, scores 0.787 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.036 in the original and +0.022 after conversion — 60 % of the delta retained.
Quality. Mean predicted overall quality across the segments went 2.90 → 3.14 (+0.24) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This is not a strict-rule trajectory. It comes from rescue rule S1, which exists because the strict rule returns nothing at all for Sadness.
The raw scorer output across the chain is 0.13, 0.26, 0.56. In the first clip the scorer found only a trace of Sadness (0.13); by the last it is at 0.56.
On the corpus-wide percentile scale those become 0.88, 0.90, 0.93 — a total move of +0.05.
That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Sadness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.05 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.
Re-ranked among the Sadness-bearing clips only — which is exactly what this rule does — the same clips read 0.08, 0.18, 0.42, a move of +0.35, which is usable again.
What this rule changes: Positives-only renormalisation. Throw away the huge block of clips that score zero -- 'no sadness' is not a degree of sadness -- and rank-normalise only the clips that actually carry the emotion. The gap disappears, and the ordinary 0.25 / 0.25 test then works normally.
What it costs: The number stops being comparable with other emotions: 'moved 0.25' now means a quarter of the sadness-BEARING subpopulation, not a quarter of the corpus. The chains also come from a small, self-selected slice of the data.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was -0.037 before conversion and 0.391 after — it rose by 0.428. Neighbour-to-neighbour the worst pair went 0.109 → 0.441. (The earlier render, with segment 1 left raw, scores 0.169 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.047 in the original and +0.507 after conversion — 1070 % of the delta retained, i.e. the move came out slightly larger after conversion than before.
Quality. Mean predicted overall quality across the segments went 2.67 → 2.85 (+0.18) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This is not a strict-rule trajectory. It comes from rescue rule S1, which exists because the strict rule returns nothing at all for Sadness.
The raw scorer output across the chain is 0.36, 0.45, 0.68. In the first clip the scorer found only a trace of Sadness (0.36); by the last it is at 0.68.
On the corpus-wide percentile scale those become 0.91, 0.92, 0.95 — a total move of +0.04.
That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Sadness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.04 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.
Re-ranked among the Sadness-bearing clips only — which is exactly what this rule does — the same clips read 0.25, 0.33, 0.54, a move of +0.28, which is usable again.
What this rule changes: Positives-only renormalisation. Throw away the huge block of clips that score zero -- 'no sadness' is not a degree of sadness -- and rank-normalise only the clips that actually carry the emotion. The gap disappears, and the ordinary 0.25 / 0.25 test then works normally.
What it costs: The number stops being comparable with other emotions: 'moved 0.25' now means a quarter of the sadness-BEARING subpopulation, not a quarter of the corpus. The chains also come from a small, self-selected slice of the data.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.258 before conversion and 0.506 after — it rose by 0.249. Neighbour-to-neighbour the worst pair went 0.274 → 0.506. (The earlier render, with segment 1 left raw, scores 0.473 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.039 in the original and +0.506 after conversion — 1289 % of the delta retained, i.e. the move came out slightly larger after conversion than before.
Quality. Mean predicted overall quality across the segments went 2.91 → 3.21 (+0.30) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This is not a strict-rule trajectory. It comes from rescue rule S1, which exists because the strict rule returns nothing at all for Sadness.
The raw scorer output across the chain is 0.11, 0.26, 0.51. In the first clip the scorer found only a trace of Sadness (0.11); by the last it is at 0.51.
On the corpus-wide percentile scale those become 0.88, 0.90, 0.92 — a total move of +0.04.
That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Sadness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.04 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.
Re-ranked among the Sadness-bearing clips only — which is exactly what this rule does — the same clips read 0.06, 0.18, 0.38, a move of +0.32, which is usable again.
What this rule changes: Positives-only renormalisation. Throw away the huge block of clips that score zero -- 'no sadness' is not a degree of sadness -- and rank-normalise only the clips that actually carry the emotion. The gap disappears, and the ordinary 0.25 / 0.25 test then works normally.
What it costs: The number stops being comparable with other emotions: 'moved 0.25' now means a quarter of the sadness-BEARING subpopulation, not a quarter of the corpus. The chains also come from a small, self-selected slice of the data.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.794 before conversion and 0.860 after — it rose by 0.066. Neighbour-to-neighbour the worst pair went 0.794 → 0.860. (The earlier render, with segment 1 left raw, scores 0.809 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.043 in the original and +0.006 after conversion — 14 % of the delta retained, so a meaningful part of the trajectory was flattened.
Quality. Mean predicted overall quality across the segments went 3.05 → 3.26 (+0.21) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This is not a strict-rule trajectory. It comes from rescue rule S1, which exists because the strict rule returns nothing at all for Sadness.
The raw scorer output across the chain is 0.67, 0.75, 1.12. In the first clip the scorer already found some Sadness here (0.67); by the last it is at 1.12.
On the corpus-wide percentile scale those become 0.94, 0.95, 0.98 — a total move of +0.04.
That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Sadness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.04 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.
Re-ranked among the Sadness-bearing clips only — which is exactly what this rule does — the same clips read 0.53, 0.61, 0.85, a move of +0.32, which is usable again.
What this rule changes: Positives-only renormalisation. Throw away the huge block of clips that score zero -- 'no sadness' is not a degree of sadness -- and rank-normalise only the clips that actually carry the emotion. The gap disappears, and the ordinary 0.25 / 0.25 test then works normally.
What it costs: The number stops being comparable with other emotions: 'moved 0.25' now means a quarter of the sadness-BEARING subpopulation, not a quarter of the corpus. The chains also come from a small, self-selected slice of the data.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.626 before conversion and 0.656 after — it rose by 0.030. Neighbour-to-neighbour the worst pair went 0.726 → 0.772. (The earlier render, with segment 1 left raw, scores 0.657 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.039 in the original and +0.017 after conversion — 45 % of the delta retained, so a meaningful part of the trajectory was flattened.
Quality. Mean predicted overall quality across the segments went 2.82 → 3.02 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This is not a strict-rule trajectory. It comes from rescue rule S1, which exists because the strict rule returns nothing at all for Sadness.
The raw scorer output across the chain is 0.28, 0.54, 0.79. In the first clip the scorer found only a trace of Sadness (0.28); by the last it is at 0.79.
On the corpus-wide percentile scale those become 0.90, 0.93, 0.96 — a total move of +0.06.
That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Sadness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.06 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.
Re-ranked among the Sadness-bearing clips only — which is exactly what this rule does — the same clips read 0.19, 0.41, 0.64, a move of +0.45, which is usable again.
What this rule changes: Positives-only renormalisation. Throw away the huge block of clips that score zero -- 'no sadness' is not a degree of sadness -- and rank-normalise only the clips that actually carry the emotion. The gap disappears, and the ordinary 0.25 / 0.25 test then works normally.
What it costs: The number stops being comparable with other emotions: 'moved 0.25' now means a quarter of the sadness-BEARING subpopulation, not a quarter of the corpus. The chains also come from a small, self-selected slice of the data.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.668 before conversion and 0.796 after — it rose by 0.128. Neighbour-to-neighbour the worst pair went 0.668 → 0.795. (The earlier render, with segment 1 left raw, scores 0.709 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.061 in the original and +0.008 after conversion — 12 % of the delta retained, so a meaningful part of the trajectory was flattened.
Quality. Mean predicted overall quality across the segments went 2.89 → 3.06 (+0.17) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This is not a strict-rule trajectory. It comes from rescue rule S1, which exists because the strict rule returns nothing at all for Sadness.
The raw scorer output across the chain is 0.09, 0.39, 0.62. In the first clip the scorer found only a trace of Sadness (0.09); by the last it is at 0.62.
On the corpus-wide percentile scale those become 0.88, 0.91, 0.94 — a total move of +0.06.
That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Sadness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.06 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.
Re-ranked among the Sadness-bearing clips only — which is exactly what this rule does — the same clips read 0.04, 0.28, 0.48, a move of +0.44, which is usable again.
What this rule changes: Positives-only renormalisation. Throw away the huge block of clips that score zero -- 'no sadness' is not a degree of sadness -- and rank-normalise only the clips that actually carry the emotion. The gap disappears, and the ordinary 0.25 / 0.25 test then works normally.
What it costs: The number stops being comparable with other emotions: 'moved 0.25' now means a quarter of the sadness-BEARING subpopulation, not a quarter of the corpus. The chains also come from a small, self-selected slice of the data.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.817 before conversion and 0.874 after — it rose by 0.058. Neighbour-to-neighbour the worst pair went 0.817 → 0.874. (The earlier render, with segment 1 left raw, scores 0.753 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
The emotional move did not survive. Re-scored end to end, Sadness moved +0.061 in the original and -0.441 after conversion — it changed direction. On this chain the corrected audio is not an improvement.
Quality. Mean predicted overall quality across the segments went 2.78 → 3.21 (+0.43) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This is not a strict-rule trajectory. It comes from rescue rule S1, which exists because the strict rule returns nothing at all for Sadness.
The raw scorer output across the chain is 0.72, 0.82, 1.20. In the first clip the scorer already found some Sadness here (0.72); by the last it is at 1.20.
On the corpus-wide percentile scale those become 0.95, 0.96, 0.99 — a total move of +0.04.
That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Sadness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.04 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.
Re-ranked among the Sadness-bearing clips only — which is exactly what this rule does — the same clips read 0.58, 0.67, 0.88, a move of +0.30, which is usable again.
What this rule changes: Positives-only renormalisation. Throw away the huge block of clips that score zero -- 'no sadness' is not a degree of sadness -- and rank-normalise only the clips that actually carry the emotion. The gap disappears, and the ordinary 0.25 / 0.25 test then works normally.
What it costs: The number stops being comparable with other emotions: 'moved 0.25' now means a quarter of the sadness-BEARING subpopulation, not a quarter of the corpus. The chains also come from a small, self-selected slice of the data.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.152 before conversion and 0.777 after — it rose by 0.625. Neighbour-to-neighbour the worst pair went 0.187 → 0.777. (The earlier render, with segment 1 left raw, scores 0.677 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.036 in the original and +0.000 after conversion — 0 % of the delta retained, so a meaningful part of the trajectory was flattened.
Quality. Mean predicted overall quality across the segments went 2.95 → 3.07 (+0.12) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This is not a strict-rule trajectory. It comes from rescue rule S1, which exists because the strict rule returns nothing at all for Sadness.
The raw scorer output across the chain is 0.44, 0.58, 0.78. In the first clip the scorer found only a trace of Sadness (0.44); by the last it is at 0.78.
On the corpus-wide percentile scale those become 0.92, 0.93, 0.96 — a total move of +0.04.
That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Sadness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.04 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.
Re-ranked among the Sadness-bearing clips only — which is exactly what this rule does — the same clips read 0.32, 0.45, 0.63, a move of +0.31, which is usable again.
What this rule changes: Positives-only renormalisation. Throw away the huge block of clips that score zero -- 'no sadness' is not a degree of sadness -- and rank-normalise only the clips that actually carry the emotion. The gap disappears, and the ordinary 0.25 / 0.25 test then works normally.
What it costs: The number stops being comparable with other emotions: 'moved 0.25' now means a quarter of the sadness-BEARING subpopulation, not a quarter of the corpus. The chains also come from a small, self-selected slice of the data.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.692 before conversion and 0.770 after — it rose by 0.078. Neighbour-to-neighbour the worst pair went 0.720 → 0.753. (The earlier render, with segment 1 left raw, scores 0.592 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
The emotional move did not survive. Re-scored end to end, Sadness moved +0.041 in the original and -0.526 after conversion — it changed direction. On this chain the corrected audio is not an improvement.
Quality. Mean predicted overall quality across the segments went 2.73 → 3.05 (+0.32) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This is not a strict-rule trajectory. It comes from rescue rule S1, which exists because the strict rule returns nothing at all for Sadness.
The raw scorer output across the chain is 0.54, 0.78, 0.95. In the first clip the scorer already found some Sadness here (0.54); by the last it is at 0.95.
On the corpus-wide percentile scale those become 0.93, 0.96, 0.97 — a total move of +0.04.
That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Sadness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.04 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.
Re-ranked among the Sadness-bearing clips only — which is exactly what this rule does — the same clips read 0.41, 0.63, 0.76, a move of +0.35, which is usable again.
What this rule changes: Positives-only renormalisation. Throw away the huge block of clips that score zero -- 'no sadness' is not a degree of sadness -- and rank-normalise only the clips that actually carry the emotion. The gap disappears, and the ordinary 0.25 / 0.25 test then works normally.
What it costs: The number stops being comparable with other emotions: 'moved 0.25' now means a quarter of the sadness-BEARING subpopulation, not a quarter of the corpus. The chains also come from a small, self-selected slice of the data.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.683 before conversion and 0.635 after — it fell by 0.048. Neighbour-to-neighbour the worst pair went 0.774 → 0.686. (The earlier render, with segment 1 left raw, scores 0.566 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
The emotional move did not survive. Re-scored end to end, Sadness moved +0.045 in the original and -0.549 after conversion — it changed direction. On this chain the corrected audio is not an improvement.
Quality. Mean predicted overall quality across the segments went 2.87 → 2.94 (+0.07) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
This is not a strict-rule trajectory. It comes from rescue rule S1, which exists because the strict rule returns nothing at all for Sadness.
The raw scorer output across the chain is 0.41, 0.61, 0.74. In the first clip the scorer found only a trace of Sadness (0.41); by the last it is at 0.74.
On the corpus-wide percentile scale those become 0.91, 0.94, 0.95 — a total move of +0.04.
That is why the strict rule cannot build this chain — but for the opposite reason to the jump case. Every clip here already carries some Sadness, so they all land in the crowded top tenth of the corpus-wide ranking, where roughly 90 % of clips scoring zero have consumed everything below. Measured that way the chain moves only 0.04 in total, far short of the 0.25 the strict rule demands — even though the raw scores clearly rise.
Re-ranked among the Sadness-bearing clips only — which is exactly what this rule does — the same clips read 0.29, 0.47, 0.60, a move of +0.31, which is usable again.
What this rule changes: Positives-only renormalisation. Throw away the huge block of clips that score zero -- 'no sadness' is not a degree of sadness -- and rank-normalise only the clips that actually carry the emotion. The gap disappears, and the ordinary 0.25 / 0.25 test then works normally.
What it costs: The number stops being comparable with other emotions: 'moved 0.25' now means a quarter of the sadness-BEARING subpopulation, not a quarter of the corpus. The chains also come from a small, self-selected slice of the data.
Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.
Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.077 before conversion and 0.781 after — it rose by 0.705. Neighbour-to-neighbour the worst pair went -0.037 → 0.563. (The earlier render, with segment 1 left raw, scores 0.666 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sadness moved +0.041 in the original and +0.488 after conversion — 1189 % of the delta retained, i.e. the move came out slightly larger after conversion than before.
Quality. Mean predicted overall quality across the segments went 2.90 → 3.26 (+0.35) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.