k-B1-k2 — voice-corrected

B1 at chain length k=2, all corpora, at the mining floor.

This is the voice-conversion-corrected version of this tier. The original, uncorrected page is still there and unchanged: t_k-B1-k2.html. Every card below carries both renders so you can switch between them without leaving the page. What was done, and what it measured →
20chains converted
20segments re-voiced
0.841 → 0.854median worst-to-anchor identity cosine
80 %median emotion delta retained
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken.

(underlined, plain · delivery, style) — the tag before the words. Emotions first, then how it is delivered. Underlined descriptors are the ones that change across this chain — anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words are a real non-speech sound, printed where it happens.

The full generated caption for any clip is under “full caption & clip details”. Its perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.

The tags and captions describe the ORIGINAL clips — they are the corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the converted audio are the ones printed in each card's own paragraph and chips.
Jealousy and Envy(unconstrained axis: Disappointment)identity +0.45 emotion 57 %   k-B1-k2 · #1

This chain comes from the one-sided rule: only Jealousy and Envy had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Jealousy and Envy clearly present — 0.75, higher than 75 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.24.

Nothing was asked of the other axis, and in fact Disappointment drifts down from 0.99 to 0.84 (-0.15), which the rule did not require.

It takes 2 clips to get there. Clip to clip the moves are +0.24 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.37 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.37 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.37, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 43 s · de · podcast

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.419 before conversion and 0.872 after — it rose by 0.453. Neighbour-to-neighbour the worst pair went 0.419 → 0.872. (The earlier render, with segment 1 left raw, scores 0.804 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Jealousy and Envy moved +0.234 in the original and +0.134 after conversion — 57 % of the delta retained. On the other named axis, Disappointment, -0.147 became -0.009.

Quality. Mean predicted overall quality across the segments went 2.48 → 3.40 (+0.92) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.419 → 0.872 +0.453identity cos neighbours 0.419 → 0.872d_b rescored +0.234 → +0.134d_a rescored -0.147 → -0.009d_a mined -0.149d_b mined 0.242min_cos_consec (site) 0.3671min_cos_anchor (site) 0.3671dataset podcastlang despeaker 580403total 42.4schain gain +2.6 dBseam step 0.1 dBcrossfades 100 ms
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: a middle-aged masculine voice · neutral-toned, balanced body, average recording, quiet background, very low-energy, neutral tension, somewhat unclear, moderate pitch range
(disappointment, bitterness, shame · measured, fairly steady, frequent disfluency, monologue) (ahem) Der Bundesrat kann im Alleingang extreme Massnahmen fordern. Was einfach ein kleiner Umwege ist, um sagen, die Leute das lesen und denken, der Bundesrat kann es alles machen, kann man verbieten, vielleicht zu
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, measured, neutral tension, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; reads as disappointment, bitterness, shame; style: monologue, whispered; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 0.8/10; 14.3s, DE.
580403_00136224 · in -18.9 dBFS · gain -1.1 dB · podcast-00484
(jealousy and envy, hope enthusiasm optimism, elation · normal-paced, moderately variable, some disfluency, monologue) Hätte das, was man schon mal als Referendum ergriff. Ja, es bleibt spannend. Es bleibt schon dort. Was ich auch immer so ein, was wir immer extrem nerven ist, das Argument mit ja, wir müssen da schon etwas machen, aber wir werden die erneuerbare Energien, weil die nicht so schnell ausbauen für Kapazität haben, dann Elektrokarren und Wärmepumpin und all das ganze Zeug. Das ist das Hauptargument der Bürgerlichen. Aber wenn es im Parlament darum geht, Solarenergie auszubauen, Winterergie auszubauen.
full caption & clip details
A young adult masculine voice; delivery is very low-energy, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as jealousy and envy, hope enthusiasm optimism, elation; style: monologue, casual; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 9.8/10; 28.2s, DE.
580403_00141992 · in -17.9 dBFS · gain -2.1 dB · podcast-02150
Concentration(unconstrained axis: Emotional Numbness)identity +0.04 emotion 138 %   k-B1-k2 · #2

This chain comes from the one-sided rule: only Concentration had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Concentration strongly present — 0.80, higher than 80 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.20.

Nothing was asked of the other axis, and in fact Emotional Numbness drifts down from 0.83 to 0.70 (-0.12), which the rule did not require.

It takes 2 clips to get there. Clip to clip the moves are +0.20 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.86 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.86 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.86. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 32 s · en · podcast

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.829 before conversion and 0.864 after — it rose by 0.035. Neighbour-to-neighbour the worst pair went 0.829 → 0.864. (The earlier render, with segment 1 left raw, scores 0.755 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.204 in the original and +0.281 after conversion — 138 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Emotional Numbness, -0.126 became -0.271.

Quality. Mean predicted overall quality across the segments went 2.92 → 3.27 (+0.35) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.829 → 0.864 +0.035identity cos neighbours 0.829 → 0.864d_b rescored +0.204 → +0.281d_a rescored -0.126 → -0.271d_a mined -0.123d_b mined 0.203min_cos_consec (site) 0.8566min_cos_anchor (site) 0.8566dataset podcastlang enspeaker 812000total 31.2schain gain +4.6 dBseam step 0.8 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a middle-aged masculine voice · neutral-toned, fairly smooth, balanced body, average recording, quiet background, slightly relaxed, frequent disfluency, somewhat unclear
(slow, very low-energy, steady, whispered) Heads of offices at the country office level, be it UN agencies or heads of implementing agencies, charities, firms,
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, slow, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: whispered, monologue; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 0.2/10; 11.0s, EN.
812000_00102424 · in -22.8 dBFS · gain +2.8 dB · podcast-03372
(concentration, contempt, contemplation · measured, subdued, fairly steady, didactic) who are the recipients of donor funding to deliver programs around poverty alleviation, who could first engage with a product like this and through that create a greater understanding of the power, the benefits of using geospatial data as
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, contempt, contemplation; style: didactic, monologue; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 2.9/10; 20.4s, EN.
812000_00103516 · in -20.7 dBFS · gain +0.7 dB · podcast-05312
Contentment(unconstrained axis: Contemplation)identity +0.05 emotion 93 %   k-B1-k2 · #3

This chain comes from the one-sided rule: only Contentment had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Contentment clearly present — 0.75, higher than 75 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.24.

Nothing was asked of the other axis, and in fact Contemplation drifts down from 0.96 to 0.89 (-0.07), which the rule did not require.

It takes 2 clips to get there. Clip to clip the moves are +0.24 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.89 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.89 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.89. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 35 s · italian · mls

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.887 before conversion and 0.933 after — it rose by 0.046. Neighbour-to-neighbour the worst pair went 0.887 → 0.933. (The earlier render, with segment 1 left raw, scores 0.670 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contentment moved +0.242 in the original and +0.225 after conversion — 93 % of the delta retained, which is essentially all of it. On the other named axis, Contemplation, -0.067 became -0.100.

Quality. Mean predicted overall quality across the segments went 2.98 → 3.38 (+0.40) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.887 → 0.933 +0.046identity cos neighbours 0.887 → 0.933d_b rescored +0.242 → +0.225d_a rescored -0.067 → -0.100d_a mined -0.067d_b mined 0.242min_cos_consec (site) 0.8898min_cos_anchor (site) 0.8898dataset mlslang italianspeaker 10446total 34.9schain gain +1.0 dBseam step 0.6 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, average recording, quiet background, normally alert, slightly relaxed, fairly steady
(contemplation, bitterness, concentration · measured, monologue, didactic) il nostro convento laico era un cattivo esempio un tradimento fatto alla società son sue parole non le rammenta sì sì le rammento ed anche le sue risposte che le parevano di trionfo l'altro dì il conte valentino chinò la testa in atto di contrizione
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation, bitterness, concentration; style: monologue, didactic; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 2.4/10; 16.8s, ITALIAN.
10446_10415_000673 · in -24.9 dBFS · gain +4.9 dB · mls-00065
(contentment, awe, pride · normal-paced, monologue, didactic) mi parevano rispose e in questo verbo è detta ogni cosa ma infine io e lei si disputava di principii si rimaneva nelle alte regioni filosofiche una donna animosa e gentile è venuta lassù con ben altri argomenti si è presentata ed ha vinto senza combattere
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as contentment, awe, pride; style: monologue, didactic; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 1.6/10; 18.4s, ITALIAN.
10446_10415_000906 · in -24.7 dBFS · gain +4.7 dB · mls-00065
Fear(unconstrained axis: Emotional Numbness)identity −0.08 emotion 241 %   k-B1-k2 · #4

This chain comes from the one-sided rule: only Fear had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Fear strongly present — 0.76, higher than 76 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.24.

Nothing was asked of the other axis, and in fact Emotional Numbness barely moves at all, sitting near 0.82 throughout.

It takes 2 clips to get there. Clip to clip the moves are +0.24 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.93 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.93 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 10 s · en · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.927 before conversion and 0.849 after — it fell by 0.077. Neighbour-to-neighbour the worst pair went 0.927 → 0.849. (The earlier render, with segment 1 left raw, scores 0.867 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fear moved +0.238 in the original and +0.574 after conversion — 241 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Emotional Numbness, +0.040 became +0.004.

Quality. Mean predicted overall quality across the segments went 2.74 → 2.80 (+0.07) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.927 → 0.849 -0.077identity cos neighbours 0.927 → 0.849d_b rescored +0.238 → +0.574d_a rescored +0.040 → +0.004d_a mined 0.040d_b mined 0.238min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_B00036_S03170total 9.7schain gain +1.6 dBseam step 0.1 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult feminine voice · neutral-toned, slightly bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(formal, monologue) Nowadays, the issue of violence on television is often debated.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 1.2/10; 4.2s, EN.
EN_B00036_S03170_W000102 · in -21.0 dBFS · gain +1.0 dB · emolia-00935
(fear, distress, disgust · formal, monologue) Many people are concerned that the images of violent acts might cause the viewers to become more aggressive.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as fear, distress, disgust; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.5/10; 5.7s, EN.
EN_B00036_S03170_W000103 · in -21.1 dBFS · gain +1.1 dB · emolia-00935
Astonishment Surprise(unconstrained axis: Jealousy and Envy)identity +0.16 emotion 303 %   k-B1-k2 · #5

This chain comes from the one-sided rule: only Astonishment Surprise had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Astonishment Surprise strongly present — 0.77, higher than 77 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.21.

Nothing was asked of the other axis, and in fact Jealousy and Envy drifts down from 0.99 to 0.06 (-0.93), which the rule did not require.

It takes 2 clips to get there. Clip to clip the moves are +0.21 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.51 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.51 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 12 s · en · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.593 before conversion and 0.749 after — it rose by 0.156. Neighbour-to-neighbour the worst pair went 0.593 → 0.749. (The earlier render, with segment 1 left raw, scores 0.567 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Astonishment Surprise moved +0.208 in the original and +0.630 after conversion — 303 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Jealousy and Envy, -0.927 became +0.627.

Quality. Mean predicted overall quality across the segments went 2.69 → 2.93 (+0.24) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.593 → 0.749 +0.156identity cos neighbours 0.593 → 0.749d_b rescored +0.208 → +0.630d_a rescored -0.927 → +0.627d_a mined -0.927d_b mined 0.206min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_yxgfczUEaOEtotal 12.1schain gain +3.4 dBseam step 0.1 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-bright, quiet background, brisk, energised, moderately variable, wide pitch range, light breath
(jealousy and envy, embarrassment, teasing · neutral tension, little disfluency, average clarity, conversational) What would make you feel comfortable with the price? Well, no one made me feel more comfortable than my kindergarten teacher, Miss Jane.
full caption & clip details
An adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, little disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as jealousy and envy, embarrassment, teasing; style: conversational, storytelling; average recording, quiet background; genuineness 1.3/6; vocal-burst blend 1.7/10; 6.5s, EN.
EN_yxgfczUEaOE_W000131 · in -20.8 dBFS · gain +0.8 dB · emolia-00926
(astonishment surprise, interest, teasing · slightly relaxed, almost no disfluency, clear, dramatic) Try saying, show me the Carfax value. You'll get the most accurate price based on the vehicle's accident history.
full caption & clip details
An adult masculine voice; delivery is energised, brisk, slightly relaxed, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, full; clear, almost no disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as astonishment surprise, interest, teasing; style: dramatic, storytelling; good recording, quiet background; genuineness 0.8/6; vocal-burst blend 1.8/10; 5.9s, EN.
EN_yxgfczUEaOE_W000132 · in -19.1 dBFS · gain -0.9 dB · emolia-00926
Contemplation(unconstrained axis: Concentration)identity +0.01 emotion 129 %   k-B1-k2 · #6

This chain comes from the one-sided rule: only Contemplation had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Contemplation strongly present — 0.75, higher than 75 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.24.

Nothing was asked of the other axis, and in fact Concentration climbs from 0.77 to 0.99 (+0.21), which the rule did not require.

It takes 2 clips to get there. Clip to clip the moves are +0.24 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.90 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.90 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 23 s · en · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.818 before conversion and 0.823 after — it rose by 0.005. Neighbour-to-neighbour the worst pair went 0.818 → 0.823. (The earlier render, with segment 1 left raw, scores 0.690 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contemplation moved +0.235 in the original and +0.304 after conversion — 129 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Concentration, +0.213 became +0.225.

Quality. Mean predicted overall quality across the segments went 2.93 → 3.06 (+0.13) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.818 → 0.823 +0.005identity cos neighbours 0.818 → 0.823d_b rescored +0.235 → +0.304d_a rescored +0.213 → +0.225d_a mined 0.214d_b mined 0.235min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_8mdfuk55_jItotal 22.4schain gain +0.9 dBseam step 0.4 dBcrossfades 100 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed
(normal-paced, little disfluency, moderate pitch range, didactic) But the Vedantic idea is not the destruction of the individual, but its real preservation.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, formal; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.8/10; 6.1s, EN.
EN_8mdfuk55_jI_W000083 · in -15.2 dBFS · gain -4.8 dB · emolia-00617
(contemplation, concentration, awe · measured, almost no disfluency, fairly narrow pitch, authoritative) We cannot prove the individual by any other means, but by referring to the universal, by proving that this individual is really the universal. If we think of the individual as separate from everything else in the universe, it cannot stand a minute.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, fairly narrow pitch, minimal breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation, concentration, awe; style: authoritative, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.7/10; 16.6s, EN.
EN_8mdfuk55_jI_W000084 · in -14.8 dBFS · gain -5.2 dB · emolia-00617
Embarrassment(unconstrained axis: Emotional Numbness)identity +0.41 emotion 100 %   k-B1-k2 · #7

This chain comes from the one-sided rule: only Embarrassment had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Embarrassment strongly present — 0.75, higher than 75 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.24.

Nothing was asked of the other axis, and in fact Emotional Numbness drifts down from 0.97 to 0.05 (-0.92), which the rule did not require.

It takes 2 clips to get there. Clip to clip the moves are +0.24 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.05 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.05 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 18 s · en · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.198 before conversion and 0.605 after — it rose by 0.407. Neighbour-to-neighbour the worst pair went 0.198 → 0.605. (The earlier render, with segment 1 left raw, scores 0.553 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Embarrassment moved +0.245 in the original and +0.244 after conversion — 100 % of the delta retained, which is essentially all of it. On the other named axis, Emotional Numbness, -0.849 became -0.568.

Quality. Mean predicted overall quality across the segments went 2.68 → 2.95 (+0.27) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.198 → 0.605 +0.407identity cos neighbours 0.198 → 0.605d_b rescored +0.245 → +0.244d_a rescored -0.849 → -0.568d_a mined -0.921d_b mined 0.245min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_8JS9omyCrmytotal 17.6schain gain +4.0 dBseam step 0.8 dBcrossfades 150 ms
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · neutral-toned, fairly smooth, average recording, quiet background, fairly steady
(emotional numbness · normal-paced, normally alert, slightly relaxed, casual) I'm gonna go back to the main room and start, start pulling people back in to that. So we'll see you in a minute.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as emotional numbness; style: casual, conversational; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 1.5/10; 6.0s, EN.
EN_8JS9omyCrmy_W000244 · in -22.8 dBFS · gain +2.8 dB · emolia-02023
(embarrassment, shame, contemplation · measured, very low-energy, relaxed, ASMR) Well, that was (low mumble) a little bit of a bummer. It was a little too short and that's my fault. My apologies. I missed, (low mumble) uh, I believe room seven and eight. I was trying to pop into all of them.
full caption & clip details
A young adult somewhat feminine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly positive, submissive, neutral openness; reads as embarrassment, shame, contemplation; style: ASMR, monologue; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 4.2/10; 11.9s, EN.
EN_8JS9omyCrmy_W000245 · in -14.5 dBFS · gain -5.5 dB · emolia-02023
Embarrassment(unconstrained axis: Contentment)identity +0.62 emotion 82 %   k-B1-k2 · #8

This chain comes from the one-sided rule: only Embarrassment had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Embarrassment strongly present — 0.76, higher than 76 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.21.

Nothing was asked of the other axis, and in fact Contentment drifts down from 0.94 to 0.76 (-0.18), which the rule did not require.

It takes 2 clips to get there. Clip to clip the moves are +0.21 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores -0.01 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst -0.01 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (-0.01, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 34 s · en · podcast

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was -0.015 before conversion and 0.607 after — it rose by 0.622. Neighbour-to-neighbour the worst pair went -0.015 → 0.607. (The earlier render, with segment 1 left raw, scores 0.510 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Embarrassment moved +0.209 in the original and +0.172 after conversion — 82 % of the delta retained, which is most of it. On the other named axis, Contentment, -0.179 became -0.103.

Quality. Mean predicted overall quality across the segments went 2.90 → 3.28 (+0.38) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 -0.015 → 0.607 +0.622identity cos neighbours -0.015 → 0.607d_b rescored +0.209 → +0.172d_a rescored -0.179 → -0.103d_a mined -0.182d_b mined 0.212min_cos_consec (site) -0.0059min_cos_anchor (site) -0.0059dataset podcastlang enspeaker 3770total 33.9schain gain +3.2 dBseam step 0.3 dBcrossfades 150 ms
Script — 2 chunks, 2 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, slightly rough, balanced body, below-average recording, quiet background, fairly steady, somewhat unclear
(contentment, hope enthusiasm optimism, relief · normal-paced, normally alert, neutral tension, casual) could stay local if you want to. I mean, you really can do it any way you want, especially again with technology, and we've all been reminded of that in the pandemic. But (low mumble) uh, you know, my wife, she's in a corporate job, she travels a lot before pandemic, and eventually hopefully will again. But (low mumble) um uh (low mumble) but I pretty much just stay in Seattle. It's a it's a local job. So there's a lot of options.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as contentment, hope enthusiasm optimism, relief; style: casual, conversational; below-average recording, quiet background; genuineness 4.5/6; vocal-burst blend 8.1/10; 21.1s, EN.
3770_00044488 · in -21.6 dBFS · gain +1.6 dB · podcast-04156
(embarrassment, relief · measured, very low-energy, relaxed, casual) Awesome. Now, (low mumble) um (low mumble) I guess to to bring it sort of into COVID, I'll ask you a question. (low mumble) Um, how did it what was the biggest way it affected the industry?
full caption & clip details
An adult masculine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly negative, submissive, slightly guarded; reads as embarrassment, relief; style: casual, conversational; below-average recording, quiet background; mildly explicit content; genuineness 3.9/6; vocal-burst blend 2.8/10; 13.0s, EN.
3770_00046632 · in -20.2 dBFS · gain +0.2 dB · podcast-04153
Doubt(unconstrained axis: Astonishment Surprise)identity −0.01 emotion 122 %   k-B1-k2 · #9

This chain comes from the one-sided rule: only Doubt had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Doubt strongly present — 0.75, higher than 75 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.23.

Nothing was asked of the other axis, and in fact Astonishment Surprise drifts down from 1.00 to 0.03 (-0.97), which the rule did not require.

It takes 2 clips to get there. Clip to clip the moves are +0.23 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.76 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.76 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.76, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 17 s · en · podcast

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.742 before conversion and 0.729 after — it fell by 0.013. Neighbour-to-neighbour the worst pair went 0.742 → 0.729. (The earlier render, with segment 1 left raw, scores 0.544 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Doubt moved +0.233 in the original and +0.285 after conversion — 122 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Astonishment Surprise, -0.996 became -0.911.

Quality. Mean predicted overall quality across the segments went 2.79 → 3.01 (+0.23) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.742 → 0.729 -0.013identity cos neighbours 0.742 → 0.729d_b rescored +0.233 → +0.285d_a rescored -0.996 → -0.911d_a mined -0.967d_b mined 0.230min_cos_consec (site) 0.7564min_cos_anchor (site) 0.7564dataset podcastlang enspeaker 500156total 16.2schain gain +5.1 dBseam step 1.2 dBcrossfades 150 ms
Script — 2 chunks, 2 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, quiet background, measured, normally alert
(astonishment surprise, confusion, intoxication altered states of consciousness · slightly relaxed, fairly steady, slurred, casual) wow that's uh (ahem) deck that's (low mumble) uh glad i didn't open that door
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; slurred, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, neutral openness; reads as astonishment surprise, confusion, intoxication altered states of consciousness; style: casual, conversational; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 0.0/10; 6.1s, EN.
500156_00161040 · in -25.1 dBFS · gain +5.2 dB · podcast-01667
(doubt · relaxed, moderately variable, somewhat unclear, casual) and more involved than I expected so (low mumble) that's (ahem) Uh actually (ahem)
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, wide pitch range, normal breath; affect is neutral, neutral stance, neutral openness; reads as doubt; style: casual, conversational; average recording, quiet background; genuineness 5.0/6; vocal-burst blend 0.0/10; 10.3s, EN.
500156_00162936 · in -26.2 dBFS · gain +6.2 dB · podcast-01665
Sourness(unconstrained axis: Infatuation)identity +0.01 emotion 24 %   k-B1-k2 · #10

This chain comes from the one-sided rule: only Sourness had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Sourness around average — 0.51, higher than 51 % of clips in this corpus — and ends with it clearly present at 0.72, higher than 72 % of clips in this corpus. That is a total rise of 0.21.

Nothing was asked of the other axis, and in fact Infatuation drifts down from 0.78 to 0.67 (-0.11), which the rule did not require.

It takes 2 clips to get there. Clip to clip the moves are +0.21 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.90 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.90 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 10 s · zh · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.807 before conversion and 0.822 after — it rose by 0.014. Neighbour-to-neighbour the worst pair went 0.807 → 0.822. (The earlier render, with segment 1 left raw, scores 0.802 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sourness moved +0.212 in the original and +0.050 after conversion — 24 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Infatuation, -0.111 became -0.334.

Quality. Mean predicted overall quality across the segments went 2.92 → 3.01 (+0.10) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.807 → 0.822 +0.014identity cos neighbours 0.807 → 0.822d_b rescored +0.212 → +0.050d_a rescored -0.111 → -0.334d_a mined -0.111d_b mined 0.207min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00059_S06881total 9.4schain gain +2.8 dBseam step 0.5 dBcrossfades 100 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, no background noise, normally alert, slightly relaxed, no disfluency, clear
(normal-paced, fairly steady, moderate pitch range, formal) 那么你便可以获得更多的财富。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; very good recording, no background noise; genuineness 0.6/6; vocal-burst blend 1.1/10; 3.0s, ZH.
ZH_B00059_S06881_W000020 · in -20.6 dBFS · gain +0.7 dB · emolia-03863
(measured, steady, fairly narrow pitch, formal) 这是富爸爸提出的三大准则的第一条,金钱只是一种观念。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; average recording, no background noise; genuineness 0.4/6; vocal-burst blend 1.0/10; 6.6s, ZH.
ZH_B00059_S06881_W000021 · in -20.5 dBFS · gain +0.5 dB · emolia-03863
Doubt(unconstrained axis: Contemplation)identity −0.02 emotion 51 %   k-B1-k2 · #11

This chain comes from the one-sided rule: only Doubt had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Doubt clearly present — 0.73, higher than 73 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.24.

Nothing was asked of the other axis, and in fact Contemplation drifts down from 0.99 to 0.92 (-0.07), which the rule did not require.

It takes 2 clips to get there. Clip to clip the moves are +0.24 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.95 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.95 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.95), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 55 s · en · podcast

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.951 before conversion and 0.927 after — it fell by 0.024. Neighbour-to-neighbour the worst pair went 0.951 → 0.927. (The earlier render, with segment 1 left raw, scores 0.694 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Doubt moved +0.245 in the original and +0.126 after conversion — 51 % of the delta retained. On the other named axis, Contemplation, -0.069 became -0.083.

Quality. Mean predicted overall quality across the segments went 2.66 → 3.22 (+0.56) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.951 → 0.927 -0.024identity cos neighbours 0.951 → 0.927d_b rescored +0.245 → +0.126d_a rescored -0.069 → -0.083d_a mined -0.066d_b mined 0.244min_cos_consec (site) 0.9516min_cos_anchor (site) 0.9516dataset podcastlang enspeaker 55799total 54.4schain gain +2.0 dBseam step 3.0 dBcrossfades 150 ms
Script — 2 chunks, 2 with a non-speech sound
Unchanged across all 2 clips: a middle-aged masculine voice · neutral-toned, slightly rough, slightly thin, average recording, quiet background, measured, neutral tension, somewhat unclear
(contemplation, interest, infatuation · very low-energy, moderately variable, some disfluency, casual) You know, it had this tiny little way to get out of it, and it was big composite thing, it weighed like a huge amount of weight. And (ahem) uh, and so after that, I told Luigi, I said, you know, if it was me, I would just go up in a spacesuit, I would ignore all this. I said, You need a space suit anyway, why not just use that? I'd strap a couple oxygen techs to the front of a tandem rig, and that's how I would do it. So anyway, Luigi's funding fell through uh (low mumble) for it, and (ahem) uh, and he decided not to buy this capsule, which was a smart idea.
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, measured, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, slightly thin; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as contemplation, interest, infatuation; style: casual, monologue; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 9.3/10; 29.5s, EN.
55799_00170976 · in -15.5 dBFS · gain -4.5 dB · podcast-01669
(doubt, relief, contemplation · subdued, fairly steady, frequent disfluency, casual) (ahem) Uh, but about I don't know, four or five months later, after the funding had gone through, it was a dead end. I call Luigi back up and I say, Hey Luigi, would you mind if I pursued this idea that you know I had suggested to you about because I thought at the time, well, tandem rig takes 450 pounds, right? More if you want to. We tested it higher than that. I said, I don't know how much a spacesuit weighs, but you know, you know,
full caption & clip details
An adult masculine voice; delivery is subdued, measured, neutral tension, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, slightly thin; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as doubt, relief, contemplation; style: casual, conversational; average recording, quiet background; genuineness 4.7/6; vocal-burst blend 8.7/10; 25.1s, EN.
55799_00173920 · in -15.6 dBFS · gain -4.4 dB · podcast-01671
Sexual Lust(unconstrained axis: Astonishment Surprise)identity −0.03 emotion 43 %   k-B1-k2 · #12

This chain comes from the one-sided rule: only Sexual Lust had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Sexual Lust strongly present — 0.79, higher than 79 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.20.

Nothing was asked of the other axis, and in fact Astonishment Surprise drifts down from 0.84 to 0.14 (-0.70), which the rule did not require.

It takes 2 clips to get there. Clip to clip the moves are +0.20 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.92 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.92 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 37 s · zh · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.896 before conversion and 0.870 after — it fell by 0.025. Neighbour-to-neighbour the worst pair went 0.896 → 0.870. (The earlier render, with segment 1 left raw, scores 0.707 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sexual Lust moved +0.201 in the original and +0.087 after conversion — 43 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Astonishment Surprise, -0.701 became -0.503.

Quality. Mean predicted overall quality across the segments went 3.12 → 3.36 (+0.24) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.896 → 0.870 -0.025identity cos neighbours 0.896 → 0.870d_b rescored +0.201 → +0.087d_a rescored -0.701 → -0.503d_a mined -0.700d_b mined 0.201min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00064_S03630total 36.2schain gain +1.3 dBseam step 0.4 dBcrossfades 100 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · balanced body, average recording, measured, fairly steady, some disfluency, slurred, light breath
(normally alert, slightly relaxed, moderate pitch range, monologue) 但是这不是说咱们要讲的,您要听的也不是王悦的爱情故事。您要听的是什么?要听的是我这梦。我昨天晚上做的那个梦到底是怎么回事?
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; slurred, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, whispered; average recording, no background noise; genuineness 3.5/6; vocal-burst blend 6.6/10; 13.9s, ZH.
ZH_B00064_S03630_W000118 · in -18.4 dBFS · gain -1.6 dB · emolia-03912
(sexual lust, infatuation, longing · subdued, relaxed, fairly narrow pitch, whispered) 为什么昨天晚上会有人在梦里边特别提示我今儿王玉儿要自杀呀?那么咱们换另外一个话说如果没有昨天,晚上那个梦我们会四处死盯着张望吗?那很有可能我这个朋友,我这个同学现在已经上了西天了。
full caption & clip details
An elderly masculine voice; delivery is subdued, measured, relaxed, fairly steady; timbre is slightly warm, dark, slightly rough, balanced body; slurred, some disfluency, fairly narrow pitch, light breath; affect is mildly negative, neutral stance, fairly guarded; reads as sexual lust, infatuation, longing; style: whispered, narration; average recording, quiet background; genuineness 2.3/6; vocal-burst blend 8.7/10; 22.4s, ZH.
ZH_B00064_S03630_W000119 · in -18.4 dBFS · gain -1.6 dB · emolia-03912
Affection(unconstrained axis: Impatience and Irritability)identity −0.01 emotion 295 %   k-B1-k2 · #13

This chain comes from the one-sided rule: only Affection had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Affection strongly present — 0.77, higher than 77 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.21.

Nothing was asked of the other axis, and in fact Impatience and Irritability drifts down from 0.94 to 0.82 (-0.12), which the rule did not require.

It takes 2 clips to get there. Clip to clip the moves are +0.21 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.91 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.91 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.91), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 25 s · en · podcast

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.900 before conversion and 0.890 after — it fell by 0.010. Neighbour-to-neighbour the worst pair went 0.900 → 0.890. (The earlier render, with segment 1 left raw, scores 0.753 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Affection moved +0.220 in the original and +0.649 after conversion — 295 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Impatience and Irritability, -0.123 became -0.105.

Quality. Mean predicted overall quality across the segments went 2.79 → 3.14 (+0.35) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.900 → 0.890 -0.010identity cos neighbours 0.900 → 0.890d_b rescored +0.220 → +0.649d_a rescored -0.123 → -0.105d_a mined -0.123d_b mined 0.214min_cos_consec (site) 0.9092min_cos_anchor (site) 0.9092dataset podcastlang enspeaker 264301total 25.0schain gain +3.7 dBseam step 0.8 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a young adult feminine voice · neutral-toned, fairly smooth, thin, average recording, quiet background, neutral tension, moderately variable, some disfluency
(impatience and irritability, pride · brisk, energised, dramatic, playful) So there's a lot of exercises. What is the exercise?
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as impatience and irritability, pride; style: dramatic, playful; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 6.1/10; 10.3s, EN.
264301_00078988 · in -23.5 dBFS · gain +3.5 dB · podcast-05729
(affection, contentment, thankfulness gratitude · normal-paced, normally alert, casual, monologue) And we have a pizza in 20 minutes, and the tennis of our swimming in 24 hours, we have this idea that we need to have a lot magical like this pastilla and this supplement.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as affection, contentment, thankfulness gratitude; style: casual, monologue; average recording, quiet background; genuineness 4.8/6; vocal-burst blend 8.6/10; 14.9s, EN.
264301_00080012 · in -23.8 dBFS · gain +3.8 dB · podcast-05716
Affection(unconstrained axis: Malevolence Malice)identity +0.42 emotion 74 %   k-B1-k2 · #14

This chain comes from the one-sided rule: only Affection had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Affection strongly present — 0.78, higher than 78 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.21.

Nothing was asked of the other axis, and in fact Malevolence Malice drifts down from 0.91 to 0.27 (-0.64), which the rule did not require.

It takes 2 clips to get there. Clip to clip the moves are +0.21 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of -0.19 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of -0.19 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 14 s · en · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was -0.103 before conversion and 0.313 after — it rose by 0.416. Neighbour-to-neighbour the worst pair went -0.103 → 0.313. (The earlier render, with segment 1 left raw, scores 0.137 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Affection moved +0.212 in the original and +0.157 after conversion — 74 % of the delta retained, which is most of it. On the other named axis, Malevolence Malice, -0.642 became -0.571.

Quality. Mean predicted overall quality across the segments went 2.64 → 2.87 (+0.22) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 -0.103 → 0.313 +0.416identity cos neighbours -0.103 → 0.313d_b rescored +0.212 → +0.157d_a rescored -0.642 → -0.571d_a mined -0.642d_b mined 0.211min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_2uZnNFc02wctotal 13.6schain gain +2.2 dBseam step 1.4 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a middle-aged masculine voice · neutral-toned, balanced body, quiet background, normally alert
(malevolence malice · measured, slightly relaxed, fairly steady, didactic) Now listen, pay attention, you're gonna watch the video completely with a, with a sound on, okay? With a sound on. Pay attention.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, slightly dominant, slightly guarded; reads as malevolence malice; style: didactic, casual; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 0.6/10; 9.5s, EN.
EN_2uZnNFc02wc_W000159 · in -17.6 dBFS · gain -2.4 dB · emolia-01442
(affection, teasing, contentment · normal-paced, relaxed, moderately variable, casual) That's okay, Jacob. I've only been here for a little while. Is everything alright?
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as affection, teasing, contentment; style: casual, conversational; below-average recording, quiet background; genuineness 4.7/6; vocal-burst blend 2.7/10; 4.3s, EN.
EN_2uZnNFc02wc_W000164 · in -21.6 dBFS · gain +1.6 dB · emolia-01442
Doubt(unconstrained axis: Sexual Lust)identity +0.02 emotion REVERSED   k-B1-k2 · #15

This chain comes from the one-sided rule: only Doubt had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Doubt strongly present — 0.79, higher than 79 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.20.

Nothing was asked of the other axis, and in fact Sexual Lust drifts down from 0.99 to 0.85 (-0.13), which the rule did not require.

It takes 2 clips to get there. Clip to clip the moves are +0.20 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.92 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.92 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 57 s · zh · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.919 before conversion and 0.939 after — it rose by 0.020. Neighbour-to-neighbour the worst pair went 0.919 → 0.939. (The earlier render, with segment 1 left raw, scores 0.823 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

The emotional move did not survive. Re-scored end to end, Doubt moved +0.200 in the original and -0.018 after conversion — it changed direction. On this chain the corrected audio is not an improvement. On the other named axis, Sexual Lust, -0.134 became -0.006.

Quality. Mean predicted overall quality across the segments went 2.84 → 3.19 (+0.35) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.919 → 0.939 +0.020identity cos neighbours 0.919 → 0.939d_b rescored +0.200 → -0.018d_a rescored -0.134 → -0.006d_a mined -0.134d_b mined 0.201min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00074_S07838total 56.8schain gain +5.0 dBseam step 0.4 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an elderly masculine voice · neutral-toned, slightly dark, balanced body, average recording, quiet background, measured, subdued, slightly relaxed
(sexual lust, intoxication altered states of consciousness, relief · fairly narrow pitch, light breath, monologue, whispered) 甘酥香脆,不油不腻,先用花椒爆香的油,再把花生米一器倒进锅内,不停地翻炒。铁锅发出动人的噗噗响声,然后就关小火,让每一粒花生米都进了油箱,最后撒一把粗盐,洗锅,放在瓷罐子里。
full caption & clip details
An elderly masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, rough, balanced body; slurred, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as sexual lust, intoxication altered states of consciousness, relief; style: monologue, whispered; average recording, quiet background; genuineness 2.0/6; vocal-burst blend 3.2/10; 27.4s, ZH.
ZH_B00074_S07838_W000000 · in -22.7 dBFS · gain +2.7 dB · emolia-04020
(doubt, jealousy and envy, triumph · moderate pitch range, audible breath, storytelling, monologue) 腌的鱼块发出一种淡淡的粉红色,霎是诱人灯,要炸了,用鸡蛋或者面粉挂上箱子,用热热的油炸的鱼块外酥里嫩,一块块金黄的摆在那里,等要吃了。用蒜头和酱油做一个浇头,淋上去又来下酒,爷爷总是会喝得摇头晃。
full caption & clip details
An elderly masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; slurred, frequent disfluency, moderate pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; reads as doubt, jealousy and envy, triumph; style: storytelling, monologue; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 5.5/10; 29.6s, ZH.
ZH_B00074_S07838_W000001 · in -21.0 dBFS · gain +1.0 dB · emolia-04020
Shame(unconstrained axis: Concentration)identity +0.01 emotion REVERSED   k-B1-k2 · #16

This chain comes from the one-sided rule: only Shame had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Shame clearly present — 0.72, higher than 72 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.24.

Nothing was asked of the other axis, and in fact Concentration barely moves at all, sitting near 0.98 throughout.

It takes 2 clips to get there. Clip to clip the moves are +0.24 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the eurospeech clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 30 s · da · eurospeech

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.884 before conversion and 0.899 after — it rose by 0.015. Neighbour-to-neighbour the worst pair went 0.884 → 0.899. (The earlier render, with segment 1 left raw, scores 0.881 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

The emotional move did not survive. Re-scored end to end, Shame moved +0.236 in the original and -0.014 after conversion — it changed direction. On this chain the corrected audio is not an improvement. On the other named axis, Concentration, -0.042 became -0.075.

Quality. Mean predicted overall quality across the segments went 3.24 → 3.47 (+0.22) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.884 → 0.899 +0.015identity cos neighbours 0.884 → 0.899d_b rescored +0.236 → -0.014d_a rescored -0.042 → -0.075d_a mined -0.042d_b mined 0.235min_cos_consec (site) —min_cos_anchor (site) —dataset eurospeechlang daspeaker denmark_20191M036_2019-12-total 29.7schain gain +4.5 dBseam step 0.1 dBcrossfades 150 ms
Script — 2 chunks, 2 with a non-speech sound
Unchanged across all 2 clips: an adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, quiet background, normally alert, some disfluency
(concentration, contemplation, confusion · brisk, neutral tension, moderately variable, cartoonish) Men (ahem) de to ting udelukker jo ikke hinanden. Vi ønsker at fastholde de stramme krav, og vi vil gerne skærpe dem. Vi kan ikke se noget problem i, at det kræves, at man har boet i landet i 20 år eller 30 år, før man kan få statsborgerskab. Så de stramme krav skal selvfølgelig være der og gerne (ahem) strammes. Men derudover
full caption & clip details
An adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as concentration, contemplation, confusion; style: cartoonish, ranting; average recording, quiet background; genuineness 2.7/6; vocal-burst blend 2.9/10; 18.0s, DA.
denmark_20191M036_2019-12-12_1000_14272912_14290944 · in -22.5 dBFS · gain +2.5 dB · eurospeech-00376
(shame, disgust, concentration · measured, slightly relaxed, fairly steady, didactic) skal det simple antal skæres ned, fordi der skal være en eller anden proportion, i forhold til hvor mange der bor (ahem) i landet. Vi kan simpelt hen ikke acceptere, at der kan skabes
full caption & clip details
A middle-aged somewhat feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as shame, disgust, concentration; style: didactic, monologue; average recording, quiet background; genuineness 2.6/6; vocal-burst blend 0.9/10; 11.9s, DA.
denmark_20191M036_2019-12-12_1000_14290944_14302880 · in -26.8 dBFS · gain +6.8 dB · eurospeech-00376
Doubt(unconstrained axis: Concentration)identity −0.05 emotion 77 %   k-B1-k2 · #17

This chain comes from the one-sided rule: only Doubt had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Doubt strongly present — 0.76, higher than 76 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.24.

Nothing was asked of the other axis, and in fact Concentration drifts down from 0.94 to 0.79 (-0.15), which the rule did not require.

It takes 2 clips to get there. Clip to clip the moves are +0.24 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.96 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.96 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 29 s · en · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.882 before conversion and 0.835 after — it fell by 0.048. Neighbour-to-neighbour the worst pair went 0.882 → 0.835. (The earlier render, with segment 1 left raw, scores 0.803 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Doubt moved +0.236 in the original and +0.182 after conversion — 77 % of the delta retained, which is most of it. On the other named axis, Concentration, -0.152 became +0.034.

Quality. Mean predicted overall quality across the segments went 2.82 → 3.16 (+0.34) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.882 → 0.835 -0.048identity cos neighbours 0.882 → 0.835d_b rescored +0.236 → +0.182d_a rescored -0.152 → +0.034d_a mined -0.152d_b mined 0.236min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_YruvfgW7Tsgtotal 28.1schain gain +2.5 dBseam step 0.9 dBcrossfades 150 ms
Script — 2 chunks, 2 with a non-speech sound
Unchanged across all 2 clips: a middle-aged masculine voice · slightly cool, slightly dark, quiet background, frequent disfluency, fairly narrow pitch, audible breath
(concentration, contemplation, thankfulness gratitude · measured, subdued, slightly relaxed, monologue) The, the maximum number of days that people could rent out via Airbnb would be increased significantly. So and I think that's kind of, I mean that kind of, that's a good image, that's a good, (low mumble) uhm, image about the, uh, (low mumble) the weird power of this company.
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is slightly cool, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, contemplation, thankfulness gratitude; style: monologue; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 0.4/10; 18.1s, EN.
EN_YruvfgW7Tsg_W000049 · in -12.5 dBFS · gain -7.5 dB · emolia-02570
(doubt, confusion, helplessness · slow, very low-energy, relaxed, didactic) Why, why would the, why would a government have to accept to give a (low mumble) concession to a company
full caption & clip details
An elderly masculine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is slightly cool, slightly dark, rough, slightly thin; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as doubt, confusion, helplessness; style: didactic, monologue; below-average recording, quiet background; genuineness 1.4/6; vocal-burst blend 0.2/10; 10.2s, EN.
EN_YruvfgW7Tsg_W000050 · in -12.4 dBFS · gain -7.6 dB · emolia-02570
Impatience and Irritability(unconstrained axis: Embarrassment)identity −0.02 emotion REVERSED   k-B1-k2 · #18

This chain comes from the one-sided rule: only Impatience and Irritability had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Impatience and Irritability clearly present — 0.70, higher than 70 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.24.

Nothing was asked of the other axis, and in fact Embarrassment drifts down from 0.97 to 0.71 (-0.27), which the rule did not require.

It takes 2 clips to get there. Clip to clip the moves are +0.24 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the eurospeech clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 33 s · da · eurospeech

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.949 before conversion and 0.925 after — it fell by 0.024. Neighbour-to-neighbour the worst pair went 0.949 → 0.925. (The earlier render, with segment 1 left raw, scores 0.792 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

The emotional move did not survive. Re-scored end to end, Impatience and Irritability moved +0.242 in the original and -0.085 after conversion — it changed direction. On this chain the corrected audio is not an improvement. On the other named axis, Embarrassment, -0.269 became -0.205.

Quality. Mean predicted overall quality across the segments went 3.18 → 3.35 (+0.17) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.949 → 0.925 -0.024identity cos neighbours 0.949 → 0.925d_b rescored +0.242 → -0.085d_a rescored -0.269 → -0.205d_a mined -0.269d_b mined 0.242min_cos_consec (site) —min_cos_anchor (site) —dataset eurospeechlang daspeaker denmark_20181M037_2018-12-total 32.9schain gain +3.1 dBseam step 0.3 dBcrossfades 150 ms
Script — 2 chunks, 2 with a non-speech sound
Unchanged across all 2 clips: an adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, quiet background, normally alert, neutral tension
(embarrassment, intoxication altered states of consciousness, elation · normal-paced, wide pitch range, ranting) man (ahem) (low mumble) indgår i dem. Og jeg er helt med på, at det ikke er de offensive missioner, vi snakker om. Er det ikke korrekt forstået, at der ligesom sker den (low mumble) ændring med det her beslutningsforslag? Det er det ene spørgsmål. Det andet er, om ministeren vil sige lidt mere om, hvad det er den her (low mumble)
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as embarrassment, intoxication altered states of consciousness, elation; style: ranting; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 3.3/10; 18.1s, DA.
denmark_20181M037_2018-12-14_1000_20056976_20075104 · in -22.0 dBFS · gain +2.0 dB · eurospeech-00347
(impatience and irritability, confusion, bitterness · brisk, moderate pitch range, monologue) mission forventelig vil gå ud på i den periode, hvor vi kommer til at deltage. Det er selvfølgelig svært at sige, men hvor forventer man at den franske hangarskibsgruppe cirka vil være henne? Hvad vil det typisk være for nogle trusler og opgaver, (low mumble) man er særlig fokuseret på? Det synes jeg kunne være rart at høre lidt mere om.
full caption & clip details
An adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as impatience and irritability, confusion, bitterness; style: monologue; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 5.8/10; 15.0s, DA.
denmark_20181M037_2018-12-14_1000_20075104_20090080 · in -22.7 dBFS · gain +2.7 dB · eurospeech-00347
Thankfulness Gratitude(unconstrained axis: Contemplation)identity +0.01 emotion 127 %   k-B1-k2 · #19

This chain comes from the one-sided rule: only Thankfulness Gratitude had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Thankfulness Gratitude strongly present — 0.77, higher than 77 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.21.

Nothing was asked of the other axis, and in fact Contemplation drifts down from 0.93 to 0.75 (-0.18), which the rule did not require.

It takes 2 clips to get there. Clip to clip the moves are +0.21 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.90 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.90 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.90. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 34 s · en · podcast

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.852 before conversion and 0.859 after — it rose by 0.006. Neighbour-to-neighbour the worst pair went 0.852 → 0.859. (The earlier render, with segment 1 left raw, scores 0.653 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Thankfulness Gratitude moved +0.203 in the original and +0.257 after conversion — 127 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Contemplation, -0.182 became -0.211.

Quality. Mean predicted overall quality across the segments went 2.93 → 3.16 (+0.23) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.852 → 0.859 +0.006identity cos neighbours 0.852 → 0.859d_b rescored +0.203 → +0.257d_a rescored -0.182 → -0.211d_a mined -0.184d_b mined 0.208min_cos_consec (site) 0.8954min_cos_anchor (site) 0.8954dataset podcastlang enspeaker 970225total 33.6schain gain +4.3 dBseam step 0.1 dBcrossfades 150 ms
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, quiet background, normal-paced, normally alert
(contemplation, distress, helplessness · slightly relaxed, fairly steady, casual, monologue) what you guys will be having is something much more different, which is we were talking about social and relationship stuff, and you guys will be severely handicapped because of the social distancing. You guys won't be able to sort of like well, I mean, there's no s dining in at places now, so that that's like one way that people do just like socialize, and it's gonna be hard to
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as contemplation, distress, helplessness; style: casual, monologue; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 10.0/10; 20.9s, EN.
970225_00144616 · in -22.5 dBFS · gain +2.5 dB · podcast-02396
(thankfulness gratitude, affection, fatigue exhaustion · neutral tension, moderately variable, casual, conversational) it's unenviable and I feel really bad for you guys, but at the same time, it's just (low mumble) um you guys are gonna have a different set of problems and you know, you just gonna have to try and prepare for that in the best way possible. Listening to this podcast might help with that. (ahem)
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as thankfulness gratitude, affection, fatigue exhaustion; style: casual, conversational; average recording, quiet background; genuineness 5.0/6; vocal-burst blend 7.3/10; 12.9s, EN.
970225_00146707 · in -22.9 dBFS · gain +2.9 dB · podcast-02404
Infatuation(unconstrained axis: Amusement)identity +0.11 emotion 13 %   k-B1-k2 · #20

This chain comes from the one-sided rule: only Infatuation had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Infatuation clearly present — 0.67, higher than 67 % of clips in this corpus — and ends with it at the very top of the corpus at 0.91, higher than 91 % of clips in this corpus. That is a total rise of 0.23.

Nothing was asked of the other axis, and in fact Amusement drifts down from 0.99 to 0.71 (-0.28), which the rule did not require.

It takes 2 clips to get there. Clip to clip the moves are +0.23 — a single step, so there is no internal shape to speak of.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.68 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.68 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.68, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 12 s · en · podcast

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.676 before conversion and 0.783 after — it rose by 0.107. Neighbour-to-neighbour the worst pair went 0.676 → 0.783. (The earlier render, with segment 1 left raw, scores 0.739 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Infatuation moved +0.233 in the original and +0.030 after conversion — 13 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Amusement, -0.273 became -0.234.

Quality. Mean predicted overall quality across the segments went 2.62 → 2.82 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2identity cos to seg 1 0.676 → 0.783 +0.107identity cos neighbours 0.676 → 0.783d_b rescored +0.233 → +0.030d_a rescored -0.273 → -0.234d_a mined -0.284d_b mined 0.233min_cos_consec (site) 0.6758min_cos_anchor (site) 0.6758dataset podcastlang enspeaker 157642total 11.5schain gain +4.0 dBseam step 1.3 dBcrossfades 150 ms
Script — 2 chunks, 2 with a non-speech sound
Unchanged across all 2 clips: a young adult feminine voice · slightly bright, fairly smooth, balanced body, average recording, moderately variable, average clarity
(amusement, teasing, contempt · brisk, energised, neutral tension, conversational) (childlike giggle) It's usually the adults. If the kids are the problem, it's usually because of adults.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as amusement, teasing, contempt; style: conversational, casual; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 1.7/10; 5.8s, EN.
157642_00739836 · in -19.4 dBFS · gain -0.6 dB · podcast-05567
(infatuation · normal-paced, normally alert, relaxed, casual) So I feel like (ahem) uh anything that (low mumble) um punishes and gives equally, like the Yule
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as infatuation; style: casual, conversational; average recording, no background noise; genuineness 3.9/6; vocal-burst blend 2.9/10; 5.9s, EN.
157642_00740408 · in -20.7 dBFS · gain +0.7 dB · podcast-04748