k-B1-k5 — voice-corrected

B1 at chain length k=5, all corpora, at the mining floor.

This is the voice-conversion-corrected version of this tier. The original, uncorrected page is still there and unchanged: t_k-B1-k5.html. Every card below carries both renders so you can switch between them without leaving the page. What was done, and what it measured →
20chains converted
80segments re-voiced
0.664 → 0.745median worst-to-anchor identity cosine
93 %median emotion delta retained
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken.

(underlined, plain · delivery, style) — the tag before the words. Emotions first, then how it is delivered. Underlined descriptors are the ones that change across this chain — anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words are a real non-speech sound, printed where it happens.

The full generated caption for any clip is under “full caption & clip details”. Its perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.

The tags and captions describe the ORIGINAL clips — they are the corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the converted audio are the ones printed in each card's own paragraph and chips.
Interest(unconstrained axis: Thankfulness Gratitude)identity −0.09 emotion 92 %   k-B1-k5 · #1

This chain comes from the one-sided rule: only Interest had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Interest below average — 0.37, lower than 63 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.59.

Nothing was asked of the other axis, and in fact Thankfulness Gratitude drifts down from 1.00 to 0.85 (-0.15), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are +0.22, then -0.04, then +0.17, then +0.24 — not a clean run: step 2 moves back the other way by 0.04 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 50 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.436 before conversion and 0.344 after — it fell by 0.093. Neighbour-to-neighbour the worst pair went 0.692 → 0.344. (The earlier render, with segment 1 left raw, scores 0.431 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Interest moved +0.590 in the original and +0.540 after conversion — 92 % of the delta retained, which is essentially all of it. On the other named axis, Thankfulness Gratitude, -0.150 became -0.040.

Quality. Mean predicted overall quality across the segments went 2.39 → 2.83 (+0.43) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.436 → 0.344 -0.093identity cos neighbours 0.692 → 0.344d_b rescored +0.590 → +0.540d_a rescored -0.150 → -0.040d_a mined -0.150d_b mined 0.589min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_VuotBan2PMktotal 48.8schain gain +5.2 dBseam step 1.1 dBcrossfades 100/150/150/150 ms
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · slightly dark, slightly thin, average recording, quiet background, fairly steady, somewhat unclear
(thankfulness gratitude, affection · measured, normally alert, slightly relaxed, casual) This afternoon, I get to work with the team from Campwood Charities in support of
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, slightly thin; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as thankfulness gratitude, affection; style: casual, monologue; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 2.8/10; 4.7s, EN.
EN_VuotBan2PMk_W000015 · in -19.8 dBFS · gain -0.2 dB · emolia-02005
(affection, hope enthusiasm optimism, contentment · slow, very low-energy, relaxed, whispered) (ahem) Uh, their team and encouraging (ahem) each other, communicating with each other.
full caption & clip details
A middle-aged somewhat feminine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is slightly warm, slightly dark, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly positive, submissive, slightly vulnerable; reads as affection, hope enthusiasm optimism, contentment; style: whispered, monologue; average recording, quiet background; genuineness 2.7/6; vocal-burst blend 3.8/10; 6.1s, EN.
EN_VuotBan2PMk_W000016 · in -16.0 dBFS · gain -4.0 dB · emolia-02005
(contemplation, doubt, pain · slow, very low-energy, relaxed, monologue) How they really do the big work and the high performing work that they do. So I'm looking forward to that.
full caption & clip details
A middle-aged feminine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is slightly cool, slightly dark, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, submissive, slightly vulnerable; reads as contemplation, doubt, pain; style: monologue, ASMR; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 3.1/10; 7.7s, EN.
EN_VuotBan2PMk_W000017 · in -18.7 dBFS · gain -1.3 dB · emolia-02005
(contentment, pride, affection · normal-paced, very low-energy, slightly relaxed, casual) We had a good visit this week from (low mumble) our partners at Cutsies Mutual. (ahem) Jackie and Morgan and Sarah from Austin and Monica joined us from their Lubbock office and we got to visit about
full caption & clip details
A middle-aged somewhat feminine voice; delivery is very low-energy, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly positive, slightly submissive, neutral openness; reads as contentment, pride, affection; style: casual, didactic; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 3.8/10; 15.2s, EN.
EN_VuotBan2PMk_W000018 · in -16.5 dBFS · gain -3.5 dB · emolia-02005
(interest, pride, contentment · slow, very low-energy, relaxed, ASMR) The great work that they're doing and the value that they place on their employees being involved and engaged in the community and some of the things that are working, some of the things that are challenging, we (low mumble) were able to have a partnership with them last fall
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly positive, slightly submissive, neutral openness; reads as interest, pride, contentment; style: ASMR, whispered; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 3.4/10; 15.9s, EN.
EN_VuotBan2PMk_W000019 · in -18.1 dBFS · gain -1.9 dB · emolia-02005
Longing(unconstrained axis: Infatuation)identity +0.15 emotion 133 %   k-B1-k5 · #2

This chain comes from the one-sided rule: only Longing had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Longing clearly present — 0.61, higher than 61 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.36.

Nothing was asked of the other axis, and in fact Infatuation drifts down from 0.99 to 0.89 (-0.10), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are +0.06, then +0.01, then +0.12, then +0.16 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.62 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.70 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.62, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 71 s · en · podcast

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.598 before conversion and 0.748 after — it rose by 0.150. Neighbour-to-neighbour the worst pair went 0.707 → 0.870. (The earlier render, with segment 1 left raw, scores 0.591 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Longing moved +0.356 in the original and +0.473 after conversion — 133 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Infatuation, -0.093 became -0.040.

Quality. Mean predicted overall quality across the segments went 2.83 → 3.15 (+0.31) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.598 → 0.748 +0.150identity cos neighbours 0.707 → 0.870d_b rescored +0.356 → +0.473d_a rescored -0.093 → -0.040d_a mined -0.099d_b mined 0.357min_cos_consec (site) 0.7050min_cos_anchor (site) 0.6203dataset podcastlang enspeaker 609785total 69.4schain gain +2.2 dBseam step 4.5 dBcrossfades 100/150/150/100 ms
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · slightly bright, fairly smooth, average recording, neutral tension, moderately variable, wide pitch range
(infatuation, pleasure ecstasy, intoxication altered states of consciousness · normal-paced, normally alert, frequent disfluency, casual) exonically it sounds is it's very very experimental. I think it's very humble, and I think that the approach that he's taking to music is just refreshing. Big time. All right. So I mean, I have my list, but I think my top is control
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, frequent disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as infatuation, pleasure ecstasy, intoxication altered states of consciousness; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.0/6; vocal-burst blend 3.5/10; 17.2s, EN.
609785_00332520 · in -13.7 dBFS · gain -6.3 dB · podcast-02538
(sexual lust, infatuation, pleasure ecstasy · brisk, energised, some disfluency, casual) by Scissor. I think she did her shit. And if you know me, I'm extremely biased when it comes to scissors. She's been one of my favorite artists since like 2013, 2014. I remember trying to put my friends on there. Yeah, she's cool.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as sexual lust, infatuation, pleasure ecstasy; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 4.5/6; vocal-burst blend 5.8/10; 14.0s, EN.
609785_00334240 · in -15.6 dBFS · gain -4.4 dB · podcast-02549
(doubt, jealousy and envy, impatience and irritability · measured, energised, frequent disfluency, casual) nah. But (low mumble) um, like just to see her progress in such a short, I mean, she's been in it for a minute, but in such a short amount of time, like she's been consistently putting out good work, but it seems like in just this past year, she has blown up
full caption & clip details
A young adult feminine voice; delivery is energised, measured, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, slightly thin; average clarity, frequent disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as doubt, jealousy and envy, impatience and irritability; style: casual, dramatic; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 4.6/10; 15.5s, EN.
609785_00336736 · in -13.6 dBFS · gain -6.4 dB · podcast-01622
(elation, pleasure ecstasy, hope enthusiasm optimism · brisk, energised, some disfluency, casual) And it's just it's just amazing seeing her grow as an artist, as a performer, as a woman, just seeing like you know, her interviews and seeing how her mindset has changed, how she's matured inside and outside of the music industry. Like,
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as elation, pleasure ecstasy, hope enthusiasm optimism; style: casual, storytelling; average recording, no background noise; mildly explicit content; genuineness 3.1/6; vocal-burst blend 5.7/10; 14.5s, EN.
609785_00338504 · in -13.7 dBFS · gain -6.3 dB · podcast-00243
(longing, affection, disgust · normal-paced, normally alert, some disfluency, casual) (ahem) And then to be (low mumble) um, you know, in a music collective where she's the only woman, like, and still being able to stay true to herself,
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as longing, affection, disgust; style: casual, conversational; average recording, no background noise; genuineness 4.1/6; vocal-burst blend 5.7/10; 8.8s, EN.
609785_00340048 · in -14.4 dBFS · gain -5.5 dB · podcast-01625
Disgust(unconstrained axis: Impatience and Irritability)identity +0.05 emotion 82 %   k-B1-k5 · #3

This chain comes from the one-sided rule: only Disgust had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Disgust clearly present — 0.65, higher than 65 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.32.

Nothing was asked of the other axis, and in fact Impatience and Irritability barely moves at all, sitting near 0.94 throughout.

It takes 5 clips to get there. Clip to clip the moves are +0.25, then +0.01, then -0.07, then +0.14 — not a clean run: step 3 moves back the other way by 0.07 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.80 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.80 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 40 s · zh · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.765 before conversion and 0.818 after — it rose by 0.053. Neighbour-to-neighbour the worst pair went 0.730 → 0.721. (The earlier render, with segment 1 left raw, scores 0.682 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disgust moved +0.322 in the original and +0.263 after conversion — 82 % of the delta retained, which is most of it. On the other named axis, Impatience and Irritability, -0.015 became +0.047.

Quality. Mean predicted overall quality across the segments went 2.84 → 3.11 (+0.27) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.765 → 0.818 +0.053identity cos neighbours 0.730 → 0.721d_b rescored +0.322 → +0.263d_a rescored -0.015 → +0.047d_a mined -0.015d_b mined 0.322min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00072_S07626total 38.5schain gain +2.6 dBseam step 1.4 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-bright
(impatience and irritability · normal-paced, energised, slightly relaxed, dramatic) 那个南鸡洞就可这局打一告吧,找一下,还有个远方的人。
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; very clear, almost no disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, fairly guarded; reads as impatience and irritability; style: dramatic, authoritative; average recording, no background noise; genuineness 1.9/6; vocal-burst blend 1.5/10; 5.9s, ZH.
ZH_B00072_S07626_W000011 · in -17.9 dBFS · gain -2.1 dB · emolia-03998
(anger, bitterness, malevolence malice · brisk, energised, neutral tension, dramatic) 停止,便在下定将所有的人情都发光融尽变糊之后就喷替你水里。
full caption & clip details
An adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, thin; very clear, almost no disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, guarded; reads as anger, bitterness, malevolence malice; style: dramatic, authoritative; average recording, no background noise; genuineness 1.7/6; vocal-burst blend 3.1/10; 7.0s, ZH.
ZH_B00072_S07626_W000012 · in -18.3 dBFS · gain -1.7 dB · emolia-03998
(impatience and irritability, contempt, pride · brisk, energised, slightly relaxed, authoritative) 如果我们出来卖木,就没法指午后独日。
full caption & clip details
An adult masculine voice; delivery is energised, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, fairly guarded; reads as impatience and irritability, contempt, pride; style: authoritative, dramatic; good recording, no background noise; genuineness 2.1/6; vocal-burst blend 1.7/10; 3.7s, ZH.
ZH_B00072_S07626_W000013 · in -18.5 dBFS · gain -1.5 dB · emolia-03998
(pain, impatience and irritability, teasing · normal-paced, normally alert, slightly relaxed, authoritative) 如果出力卖母,因为病后身体无力,必定献丑。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as pain, impatience and irritability, teasing; style: authoritative, didactic; average recording, no background noise; genuineness 1.9/6; vocal-burst blend 1.5/10; 4.9s, ZH.
ZH_B00072_S07626_W000014 · in -18.5 dBFS · gain -1.5 dB · emolia-03998
(disgust, bitterness, malevolence malice · normal-paced, highly aroused, neutral tension, authoritative) 请各位汉官皇家师傅多多包涵,就当你就在我这个异乡人那样想,我挤门前滚滚,反是他公云打着陆军,让毛着鹿发枪打允叫的尤勤了。
full caption & clip details
A middle-aged masculine voice; delivery is highly aroused, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, rough, balanced body; very clear, almost no disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, fairly guarded; reads as disgust, bitterness, malevolence malice; style: authoritative, cartoonish; average recording, quiet background; genuineness 1.3/6; vocal-burst blend 4.2/10; 17.7s, ZH.
ZH_B00072_S07626_W000015 · in -17.6 dBFS · gain -2.4 dB · emolia-03998
Concentration(unconstrained axis: Emotional Numbness)identity +0.09 emotion 78 %   k-B1-k5 · #4

This chain comes from the one-sided rule: only Concentration had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Concentration around average — 0.56, higher than 56 % of clips in this corpus — and ends with it at the very top of the corpus at 0.92, higher than 92 % of clips in this corpus. That is a total rise of 0.37.

Nothing was asked of the other axis, and in fact Emotional Numbness drifts down from 0.94 to 0.83 (-0.11), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are +0.16, then +0.07, then -0.09, then +0.24 — not a clean run: step 3 moves back the other way by 0.09 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 35 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.623 before conversion and 0.715 after — it rose by 0.092. Neighbour-to-neighbour the worst pair went 0.657 → 0.692. (The earlier render, with segment 1 left raw, scores 0.553 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.368 in the original and +0.286 after conversion — 78 % of the delta retained, which is most of it. On the other named axis, Emotional Numbness, -0.108 became +0.007.

Quality. Mean predicted overall quality across the segments went 2.72 → 2.94 (+0.22) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.623 → 0.715 +0.092identity cos neighbours 0.657 → 0.692d_b rescored +0.368 → +0.286d_a rescored -0.108 → +0.007d_a mined -0.108d_b mined 0.369min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN__wlvsXPnEvAtotal 34.0schain gain +2.9 dBseam step 0.5 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normal-paced, normally alert, slightly relaxed
(emotional numbness, fear · some disfluency, average clarity, casual, conversational) It was clear law enforcement is ahead by 42%, followed by a fire rescue at about 38%.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as emotional numbness, fear; style: casual, conversational; good recording, quiet background; genuineness 2.8/6; vocal-burst blend 1.3/10; 5.1s, EN.
EN__wlvsXPnEvA_W000174 · in -18.9 dBFS · gain -1.1 dB · emolia-01639
(fear · some disfluency, average clarity, casual, monologue) Emergency management at about 12% and other areas like search and rescue, (low mumble) uh, EMS starting to come into play (ahem) with the remainder.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as fear; style: casual, monologue; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 1.0/10; 8.1s, EN.
EN__wlvsXPnEvA_W000175 · in -18.2 dBFS · gain -1.8 dB · emolia-01639
(some disfluency, average clarity, casual, monologue) When we looked at the use cases, this is the response we got back. It was over 17 public safety use cases for drones.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: casual, monologue; good recording, quiet background; genuineness 3.1/6; vocal-burst blend 1.5/10; 8.1s, EN.
EN__wlvsXPnEvA_W000176 · in -17.7 dBFS · gain -2.3 dB · emolia-01639
(interest, doubt · frequent disfluency, somewhat unclear, didactic, monologue) And what is interesting about this is in addition to this, this is this list of
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as interest, doubt; style: didactic, monologue; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 0.0/10; 4.9s, EN.
EN__wlvsXPnEvA_W000177 · in -17.6 dBFS · gain -2.4 dB · emolia-01639
(concentration · some disfluency, average clarity, monologue, didactic) Use cases, you can subdivide each one of these down into three or four smaller subcategories (ahem) where we're seeing it being specifically used.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: monologue, didactic; average recording, quiet background; genuineness 1.7/6; vocal-burst blend 0.2/10; 8.6s, EN.
EN__wlvsXPnEvA_W000178 · in -18.4 dBFS · gain -1.6 dB · emolia-01639
Elation(unconstrained axis: Intoxication Altered States of Consciousness)identity +0.07 emotion 96 %   k-B1-k5 · #5

This chain comes from the one-sided rule: only Elation had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Elation clearly present — 0.73, higher than 73 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.25.

Nothing was asked of the other axis, and in fact Intoxication Altered States of Consciousness climbs from 0.83 to 0.91 (+0.08), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are +0.13, then -0.05, then +0.16, then +0.01 — not a clean run: step 2 moves back the other way by 0.05 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.84 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.84 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 43 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.672 before conversion and 0.744 after — it rose by 0.072. Neighbour-to-neighbour the worst pair went 0.780 → 0.805. (The earlier render, with segment 1 left raw, scores 0.642 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Elation moved +0.253 in the original and +0.244 after conversion — 96 % of the delta retained, which is essentially all of it. On the other named axis, Intoxication Altered States of Consciousness, +0.083 became +0.133.

Quality. Mean predicted overall quality across the segments went 2.78 → 3.06 (+0.28) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.672 → 0.744 +0.072identity cos neighbours 0.780 → 0.805d_b rescored +0.253 → +0.244d_a rescored +0.083 → +0.133d_a mined 0.083d_b mined 0.253min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_580NM9Oin1Ytotal 42.1schain gain +4.2 dBseam step 0.9 dBcrossfades 100/150/150/150 ms
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, fairly smooth, balanced body, quiet background, some disfluency, average clarity, light breath
(normal-paced, normally alert, slightly relaxed, casual) The Drupals, which I like to break the ice with. And Cogname Modelling. And...
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; no dominant emotion; style: casual, conversational; good recording, quiet background; genuineness 3.0/6; vocal-burst blend 1.6/10; 4.4s, EN.
EN_580NM9Oin1Y_W000014 · in -22.9 dBFS · gain +2.9 dB · emolia-02101
(interest, hope enthusiasm optimism, embarrassment · brisk, energised, slightly relaxed, dramatic) So I want to start by just breaking down the anatomy of content, right? So let's, let's just try and just deconstruct this and let's not talk about Drupal for a second. Let's just break some content down.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as interest, hope enthusiasm optimism, embarrassment; style: dramatic, casual; good recording, quiet background; genuineness 2.2/6; vocal-burst blend 2.3/10; 9.6s, EN.
EN_580NM9Oin1Y_W000016 · in -18.5 dBFS · gain -1.5 dB · emolia-02101
(embarrassment, sexual lust, amusement · brisk, normally alert, slightly relaxed, casual) So, I realize I've got a lapel mark, but because I've got so many slides, I have to stand here anyway, so I should have one of these clickers.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as embarrassment, sexual lust, amusement; style: casual, conversational; good recording, quiet background; genuineness 3.3/6; vocal-burst blend 4.5/10; 5.8s, EN.
EN_580NM9Oin1Y_W000017 · in -19.1 dBFS · gain -0.9 dB · emolia-02101
(elation, hope enthusiasm optimism, pleasure ecstasy · brisk, energised, slightly relaxed, casual) Anatomy of web content. Sorry, I'm gonna break this down to like about seven different things. So we've got a page that's really, really obvious. So everyone understands a page. I've got an about page on the internet. Sweet.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as elation, hope enthusiasm optimism, pleasure ecstasy; style: casual, monologue; good recording, quiet background; genuineness 2.5/6; vocal-burst blend 3.5/10; 10.4s, EN.
EN_580NM9Oin1Y_W000018 · in -19.1 dBFS · gain -0.9 dB · emolia-02101
(elation, interest, hope enthusiasm optimism · normal-paced, normally alert, neutral tension, casual) (low mumble) Uhm, so there's sort of like aspects of this, right? Like it's gonna have like a clean URL, like an SEO friendly URL. It's gonna be something that appears in the search result. You know, it's gonna have, it's gonna want to have SEO.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, slightly guarded; reads as elation, interest, hope enthusiasm optimism; style: casual, conversational; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 6.6/10; 12.6s, EN.
EN_580NM9Oin1Y_W000019 · in -18.0 dBFS · gain -2.0 dB · emolia-02101
Infatuation(unconstrained axis: Doubt)identity +0.09 emotion REVERSED   k-B1-k5 · #6

This chain comes from the one-sided rule: only Infatuation had to get where it was going, by at least 0.50. The other emotion was left completely free.

The chain starts with Infatuation below average — 0.30, lower than 70 % of clips in this corpus — and ends with it strongly present at 0.86, higher than 86 % of clips in this corpus. That is a total rise of 0.56.

Nothing was asked of the other axis, and in fact Doubt drifts down from 0.91 to 0.11 (-0.79), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are -0.05, then +0.17, then +0.21, then +0.23 — not a clean run: step 1 moves back the other way by 0.05 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 31 s · ja · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.653 before conversion and 0.739 after — it rose by 0.087. Neighbour-to-neighbour the worst pair went 0.546 → 0.654. (The earlier render, with segment 1 left raw, scores 0.581 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

The emotional move did not survive. Re-scored end to end, Infatuation moved +0.555 in the original and -0.123 after conversion — it changed direction. On this chain the corrected audio is not an improvement. On the other named axis, Doubt, -0.794 became -0.554.

Quality. Mean predicted overall quality across the segments went 2.89 → 3.14 (+0.25) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.653 → 0.739 +0.087identity cos neighbours 0.546 → 0.654d_b rescored +0.555 → -0.123d_a rescored -0.794 → -0.554d_a mined -0.794d_b mined 0.555min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang jaspeaker JA_B00004_S07042total 29.8schain gain -1.6 dBseam step 1.1 dBcrossfades 100/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, no background noise, measured, normally alert, slightly relaxed, light breath
(doubt · fairly steady, some disfluency, somewhat unclear, monologue) ということでWi-Fi メディア配信の実施例をご紹介します。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as doubt; style: monologue, narration; average recording, no background noise; genuineness 2.0/6; vocal-burst blend 3.0/10; 5.6s, JA.
JA_B00004_S07042_W000046 · in -15.1 dBFS · gain -4.9 dB · emolia-02995
(doubt · fairly steady, no disfluency, clear, formal) 画面の六行を文末に追記します。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as doubt; style: formal, authoritative; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 1.5/10; 3.4s, JA.
JA_B00004_S07042_W000047 · in -15.2 dBFS · gain -4.8 dB · emolia-02995
(steady, no disfluency, clear, formal) 画面のようにタイプして、新規ファイルを作ります。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, didactic; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.6/10; 4.0s, JA.
JA_B00004_S07042_W000048 · in -15.6 dBFS · gain -4.4 dB · emolia-02995
(thankfulness gratitude, awe · steady, no disfluency, slurred, monologue) VLC以外にも、ゲームプレイヤー、オープレイヤーライトも利用可能ですが、VLCが動作が軽快です。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; slurred, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as thankfulness gratitude, awe; style: monologue, narration; average recording, no background noise; genuineness 0.6/6; vocal-burst blend 1.5/10; 9.8s, JA.
JA_B00004_S07042_W000049 · in -16.5 dBFS · gain -3.5 dB · emolia-02995
(fairly steady, little disfluency, average clarity, didactic) パイ3でも快適で、オーバークロックは必要ありませんが、やります。画面のようにタイプして、
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, monologue; average recording, no background noise; genuineness 0.9/6; vocal-burst blend 1.4/10; 7.8s, JA.
JA_B00004_S07042_W000050 · in -16.0 dBFS · gain -4.0 dB · emolia-02995
Confusion(unconstrained axis: Anger)identity +0.03 emotion 33 %   k-B1-k5 · #7

This chain comes from the one-sided rule: only Confusion had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Confusion clearly present — 0.66, higher than 66 % of clips in this corpus — and ends with it at the very top of the corpus at 0.91, higher than 91 % of clips in this corpus. That is a total rise of 0.25.

Nothing was asked of the other axis, and in fact Anger drifts down from 0.86 to 0.01 (-0.86), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are +0.09, then -0.09, then +0.21, then +0.05 — not a clean run: step 2 moves back the other way by 0.09 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 31 s · zh · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.703 before conversion and 0.731 after — it rose by 0.028. Neighbour-to-neighbour the worst pair went 0.695 → 0.768. (The earlier render, with segment 1 left raw, scores 0.637 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Confusion moved +0.253 in the original and +0.082 after conversion — 33 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Anger, -0.857 became -0.043.

Quality. Mean predicted overall quality across the segments went 2.82 → 3.00 (+0.18) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.703 → 0.731 +0.028identity cos neighbours 0.695 → 0.768d_b rescored +0.253 → +0.082d_a rescored -0.857 → -0.043d_a mined -0.857d_b mined 0.253min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00078_S07187total 29.7schain gain +2.5 dBseam step 1.9 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, balanced body, average recording, normally alert, some disfluency
(measured, slightly relaxed, fairly steady, formal) 然后推暗自推动这个整个这个三国的一个历史的进程。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; average recording, no background noise; genuineness 2.8/6; vocal-burst blend 1.4/10; 4.5s, ZH.
ZH_B00078_S07187_W000162 · in -20.4 dBFS · gain +0.4 dB · emolia-04059
(sexual lust, intoxication altered states of consciousness, impatience and irritability · normal-paced, neutral tension, moderately variable, casual) 嗯,我爸其实喜欢玩那个,然后后来我也玩,然后我爸我玩的时候,我爸都快被气死了。
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as sexual lust, intoxication altered states of consciousness, impatience and irritability; style: casual, conversational; average recording, some background noise; mildly explicit content; genuineness 5.9/6; vocal-burst blend 2.5/10; 5.7s, ZH.
ZH_B00078_S07187_W000163 · in -20.8 dBFS · gain +0.8 dB · emolia-04059
(normal-paced, slightly relaxed, fairly steady, casual) 然后当时那个最开始选那个阶段嘛,就像什么刘备啊,他们还是比较小一个,就几个城市。
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 4.6/6; vocal-burst blend 6.9/10; 5.7s, ZH.
ZH_B00078_S07187_W000164 · in -20.7 dBFS · gain +0.7 dB · emolia-04059
(impatience and irritability, doubt, disappointment · normal-paced, neutral tension, fairly steady, casual) (low mumble) 然后我前期发展发展完了以后,我就推。然后我爸说那你就把那个张飞刘关张你给降服了,他要不降,你放了你再抓,然后没有我抓着全给砍了。
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as impatience and irritability, doubt, disappointment; style: casual, conversational; average recording, quiet background; genuineness 4.7/6; vocal-burst blend 8.3/10; 10.1s, ZH.
ZH_B00078_S07187_W000165 · in -18.9 dBFS · gain -1.1 dB · emolia-04059
(confusion · measured, slightly relaxed, fairly steady, casual) 怎么说呢?你作为一个事业,然后技术场战斗,然后有a有那个if线。
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as confusion; style: casual, conversational; average recording, quiet background; genuineness 5.2/6; vocal-burst blend 5.1/10; 4.5s, ZH.
ZH_B00078_S07187_W000166 · in -21.5 dBFS · gain +1.5 dB · emolia-04059
Emotional Numbness(unconstrained axis: Intoxication Altered States of Consciousness)identity +0.08 emotion 106 %   k-B1-k5 · #8

This chain comes from the one-sided rule: only Emotional Numbness had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Emotional Numbness clearly present — 0.60, higher than 60 % of clips in this corpus — and ends with it at the very top of the corpus at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.33.

Nothing was asked of the other axis, and in fact Intoxication Altered States of Consciousness drifts down from 0.80 to 0.40 (-0.40), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are +0.14, then +0.24, then -0.19, then +0.15 — not a clean run: step 3 moves back the other way by 0.19 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.78 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.78 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 28 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.724 before conversion and 0.809 after — it rose by 0.084. Neighbour-to-neighbour the worst pair went 0.724 → 0.809. (The earlier render, with segment 1 left raw, scores 0.750 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.332 in the original and +0.354 after conversion — 106 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Intoxication Altered States of Consciousness, -0.401 became -0.302.

Quality. Mean predicted overall quality across the segments went 2.84 → 2.96 (+0.12) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.724 → 0.809 +0.084identity cos neighbours 0.724 → 0.809d_b rescored +0.332 → +0.354d_a rescored -0.401 → -0.302d_a mined -0.401d_b mined 0.332min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_2s96uDjHP38total 26.2schain gain +0.8 dBseam step 1.1 dBcrossfades 150/100/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed
(measured, fairly steady, formal, monologue) Therefore, the proteome level shows a higher molecular complexity.
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.4/10; 4.1s, EN.
EN_2s96uDjHP38_W000005 · in -14.8 dBFS · gain -5.2 dB · emolia-01609
(pain · measured, steady, formal, monologue) The proteome is defined as the entirety of proteins expressed by a genome or by a cell tissue at a given time.
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as pain; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 7.1s, EN.
EN_2s96uDjHP38_W000006 · in -15.6 dBFS · gain -4.4 dB · emolia-01609
(emotional numbness · measured, steady, formal, monologue) Post-translational modifications, protein turnover and subcellular localization.
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: formal, monologue; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.0/10; 5.1s, EN.
EN_2s96uDjHP38_W000008 · in -16.4 dBFS · gain -3.6 dB · emolia-01609
(normal-paced, fairly steady, formal, monologue) Mass spectrometry is the standard method for proteomic analyses of complex samples.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.2/10; 5.2s, EN.
EN_2s96uDjHP38_W000009 · in -18.2 dBFS · gain -1.8 dB · emolia-01609
(emotional numbness · normal-paced, steady, formal, monologue) In the classical bottom-up approach, proteins are enzymatically digested into peptides.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 5.4s, EN.
EN_2s96uDjHP38_W000010 · in -14.4 dBFS · gain -5.6 dB · emolia-01609
Concentration(unconstrained axis: Awe)identity −0.05 emotion 94 %   k-B1-k5 · #9

This chain comes from the one-sided rule: only Concentration had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Concentration strongly present — 0.75, higher than 75 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.22.

Nothing was asked of the other axis, and in fact Awe drifts down from 0.96 to 0.47 (-0.49), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are +0.19, then -0.01, then -0.19, then +0.23 — not a clean run: step 2 moves back the other way by 0.01 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.95 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.95 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 49 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.917 before conversion and 0.866 after — it fell by 0.051. Neighbour-to-neighbour the worst pair went 0.873 → 0.855. (The earlier render, with segment 1 left raw, scores 0.567 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.219 in the original and +0.206 after conversion — 94 % of the delta retained, which is essentially all of it. On the other named axis, Awe, -0.488 became -0.485.

Quality. Mean predicted overall quality across the segments went 2.91 → 3.07 (+0.16) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.917 → 0.866 -0.051identity cos neighbours 0.873 → 0.855d_b rescored +0.219 → +0.206d_a rescored -0.488 → -0.485d_a mined -0.488d_b mined 0.219min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_jJk2kYEXAJgtotal 47.3schain gain +1.1 dBseam step 0.3 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(awe · fairly steady, formal, newsreading) Measurements of the radial velocity and proper motion of stars allows astronomers to plot the movement of these systems through the Milky Way galaxy
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as awe; style: formal, newsreading; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.0/10; 7.8s, EN.
EN_jJk2kYEXAJg_W000119 · in -14.3 dBFS · gain -5.7 dB · emolia-02339
(awe, concentration · steady, newsreading, formal) Astrometric results are the basis used to calculate the distribution of speculated dark matter in the galaxy.During the 1990s, the measurement of the stellar wobble of nearby stars was used to detect large extrasolar planets orbiting those stars.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as awe, concentration; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 13.7s, EN.
EN_jJk2kYEXAJg_W000120 · in -14.8 dBFS · gain -5.2 dB · emolia-02339
(concentration · fairly steady, newsreading, formal) Theoretical astronomers use several tools including analytical models and computational numerical simulations, each has its particular advantages
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: newsreading, formal; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.0/10; 8.7s, EN.
EN_jJk2kYEXAJg_W000122 · in -14.7 dBFS · gain -5.3 dB · emolia-02339
(fairly steady, formal, authoritative) Analytical models of a process are generally better for giving broader insight into the heart of what is going on
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.4/10; 6.0s, EN.
EN_jJk2kYEXAJg_W000123 · in -13.9 dBFS · gain -6.1 dB · emolia-02339
(concentration, emotional numbness · steady, newsreading, formal) Numerical models reveal the existence of phenomena and effects otherwise unobserved.Theorists in astronomy endeavor to create theoretical models and from the results predict observational consequences of those models.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, emotional numbness; style: newsreading, formal; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.0/10; 11.8s, EN.
EN_jJk2kYEXAJg_W000124 · in -14.4 dBFS · gain -5.6 dB · emolia-02339
Intoxication Altered States of Consciousness(unconstrained axis: Relief)identity +0.18 emotion 119 %   k-B1-k5 · #10

This chain comes from the one-sided rule: only Intoxication Altered States of Consciousness had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Intoxication Altered States of Consciousness clearly present — 0.63, higher than 63 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.32.

Nothing was asked of the other axis, and in fact Relief drifts down from 0.95 to 0.29 (-0.66), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are +0.22, then +0.13, then +0.01, then -0.03 — not a clean run: step 4 moves back the other way by 0.03 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.56 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.64 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.56, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 57 s · pl · podcast

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.212 before conversion and 0.391 after — it rose by 0.178. Neighbour-to-neighbour the worst pair went 0.212 → 0.391. (The earlier render, with segment 1 left raw, scores 0.323 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Intoxication Altered States of Consciousness moved +0.308 in the original and +0.367 after conversion — 119 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Relief, -0.663 became -0.620.

Quality. Mean predicted overall quality across the segments went 2.98 → 3.19 (+0.21) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.212 → 0.391 +0.178identity cos neighbours 0.212 → 0.391d_b rescored +0.308 → +0.367d_a rescored -0.663 → -0.620d_a mined -0.662d_b mined 0.323min_cos_consec (site) 0.6404min_cos_anchor (site) 0.5599dataset podcastlang plspeaker 863538total 56.0schain gain +2.7 dBseam step 2.9 dBcrossfades 150/100/150/150 ms
Script — 5 chunks, 4 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, quiet background
(relief, interest, teasing · normal-paced, normally alert, slightly relaxed, monologue) Tak jak też (low mumble) już wspomniałe, że Pelżem Nirwana, Alice Change i soundgrad. (ahem) Ten kawałek nie jest taki typowo grandowy, bo on jest taki brudny z jednej strony i taki pełen tych nerwów i emocji na wierzchu. A z drugiej strony jest według mnie.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as relief, interest, teasing; style: monologue, casual; good recording, quiet background; mildly explicit content; genuineness 3.2/6; vocal-burst blend 6.5/10; 14.7s, PL.
863538_00079528 · in -29.1 dBFS · gain +9.1 dB · podcast-04848
(slow, very low-energy, fully relaxed, casual) brudnawy. I wokalnie no (ahem) właśnie.
full caption & clip details
An adult masculine voice; delivery is very low-energy, slow, fully relaxed, moderately variable; timbre is neutral-toned, dark, gravelly, thin; slurred, frequent disfluency, wide pitch range, breathless; affect is mildly negative, neutral stance, neutral openness; style: casual, conversational; average recording, quiet background; no dominant emotion; genuineness 3.8/6; vocal-burst blend 0.6/10; 3.2s, PL.
863538_00081264 · in -28.2 dBFS · gain +8.2 dB · podcast-04852
(embarrassment, intoxication altered states of consciousness, sourness · fast, normally alert, slightly relaxed, casual) Bo rozmawialiśmy o tym nie tak dawno temu. Z rok temu będzie albo lepiej, nie? No
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, neutral openness; reads as embarrassment, intoxication altered states of consciousness, sourness; style: casual, conversational; average recording, quiet background; genuineness 4.8/6; vocal-burst blend 2.4/10; 4.1s, PL.
863538_00083816 · in -28.1 dBFS · gain +8.1 dB · podcast-04846
(intoxication altered states of consciousness, hope enthusiasm optimism, pride · measured, subdued, neutral tension, casual) Że jest to muzyka młoda, (ahem) (ahem) młodego pokolenia, młodego pokolenia (ahem) lat końca lat 80. i początku lat 90. Czyli jest to pokolenie X. Zdefiniowane w demografii i socjologii, które właśnie wchodzi w dorosłość. Pokolenie X, które (ahem) (ahem) nawet trochę się buntuje, ale tak naprawdę nie do końca wie (low mumble) przed czym pokolenie, które rzeczywiście trochę ma w nosie
full caption & clip details
A young adult masculine voice; delivery is subdued, measured, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, slightly submissive, slightly guarded; reads as intoxication altered states of consciousness, hope enthusiasm optimism, pride; style: casual, monologue; below-average recording, quiet background; genuineness 3.9/6; vocal-burst blend 8.2/10; 28.8s, PL.
863538_00084560 · in -26.8 dBFS · gain +6.8 dB · podcast-01441
(intoxication altered states of consciousness, pain · slow, normally alert, relaxed, casual) jest trochę zblazowane i (low mumble) a jednocześnie bardzo
full caption & clip details
An adult masculine voice; delivery is normally alert, slow, relaxed, fairly steady; timbre is neutral-toned, dark, fairly smooth, balanced body; slurred, frequent disfluency, moderate pitch range, normal breath; affect is neutral, neutral stance, neutral openness; reads as intoxication altered states of consciousness, pain; style: casual, conversational; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 0.7/10; 5.8s, PL.
863538_00087441 · in -27.2 dBFS · gain +7.2 dB · podcast-05621
Embarrassment(unconstrained axis: Interest)identity −0.03 emotion 126 %   k-B1-k5 · #11

This chain comes from the one-sided rule: only Embarrassment had to get where it was going, by at least 0.50. The other emotion was left completely free.

The chain starts with Embarrassment below average — 0.31, lower than 69 % of clips in this corpus — and ends with it at the very top of the corpus at 0.92, higher than 92 % of clips in this corpus. That is a total rise of 0.61.

Nothing was asked of the other axis, and in fact Interest drifts down from 0.94 to 0.71 (-0.23), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are +0.19, then +0.14, then +0.09, then +0.19 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.82 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.83 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.82. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 60 s · en · podcast

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.816 before conversion and 0.788 after — it fell by 0.028. Neighbour-to-neighbour the worst pair went 0.815 → 0.786. (The earlier render, with segment 1 left raw, scores 0.659 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Embarrassment moved +0.612 in the original and +0.771 after conversion — 126 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Interest, -0.228 became -0.186.

Quality. Mean predicted overall quality across the segments went 3.07 → 3.23 (+0.16) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.816 → 0.788 -0.028identity cos neighbours 0.815 → 0.786d_b rescored +0.612 → +0.771d_a rescored -0.228 → -0.186d_a mined -0.228d_b mined 0.613min_cos_consec (site) 0.8323min_cos_anchor (site) 0.8239dataset podcastlang enspeaker 870055total 58.7schain gain +2.7 dBseam step 1.1 dBcrossfades 100/150/150/100 ms
Script — 5 chunks, 4 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, balanced body, quiet background
(interest, hope enthusiasm optimism · normal-paced, normally alert, neutral tension, casual) you know, not only stories from from my book, but you know, stuff in the media that we see in weird stories. So for everyone out there in podcast land, if you guys have a story
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as interest, hope enthusiasm optimism; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 4.3/6; vocal-burst blend 7.0/10; 10.6s, EN.
870055_00212119 · in -28.7 dBFS · gain +8.7 dB · podcast-03486
(hope enthusiasm optimism, elation · measured, subdued, relaxed, casual) that (low mumble) um you see in the news uh that you want us to talk about, you want us to do some reading on and you want us to talk about, send us a link. (low mumble) Um you can send that to info
full caption & clip details
An adult masculine voice; delivery is subdued, measured, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly positive, neutral stance, neutral openness; reads as hope enthusiasm optimism, elation; style: casual, monologue; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 3.4/10; 11.0s, EN.
870055_00213176 · in -30.0 dBFS · gain +10.0 dB · podcast-03488
(normal-paced, very low-energy, slightly relaxed, casual) at T R D Tmedia dot com. That's info at TRDT Media dot com. Feel free to send it to us (low mumble) uh and we can read it. And if we feel like it's something that (low mumble) uh, you know, fits what we're doing here,
full caption & clip details
A young adult masculine voice; delivery is very low-energy, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; average recording, quiet background; genuineness 3.3/6; vocal-burst blend 3.9/10; 14.6s, EN.
870055_00214271 · in -29.8 dBFS · gain +9.8 dB · podcast-03488
(elation, hope enthusiasm optimism, thankfulness gratitude · measured, normally alert, slightly relaxed, casual) more than happy to make it into an episode and and discuss it. And (ahem) uh who knows, maybe we'll even (low mumble) uh give you a call and
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as elation, hope enthusiasm optimism, thankfulness gratitude; style: casual, monologue; good recording, quiet background; genuineness 2.7/6; vocal-burst blend 2.7/10; 7.4s, EN.
870055_00215728 · in -29.8 dBFS · gain +9.8 dB · podcast-04648
(embarrassment, contentment · measured, subdued, relaxed, casual) Yeah, if and also if you uh (low mumble) if if you if you're in the profession, if you're prior military and you've you've you've seen some action, (low mumble) um or if you know you're uh you're law enforcement, you're in forensics, (low mumble) um and you want to be on the show, same thing. Hit us up.
full caption & clip details
An adult masculine voice; delivery is subdued, measured, relaxed, fairly steady; timbre is neutral-toned, very dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly positive, submissive, neutral openness; reads as embarrassment, contentment; style: casual, monologue; average recording, quiet background; genuineness 4.6/6; vocal-burst blend 4.2/10; 15.8s, EN.
870055_00219440 · in -31.3 dBFS · gain +11.3 dB · podcast-03485
Interest(unconstrained axis: Doubt)identity −0.01 emotion 72 %   k-B1-k5 · #12

This chain comes from the one-sided rule: only Interest had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Interest below average — 0.35, lower than 65 % of clips in this corpus — and ends with it at the very top of the corpus at 0.91, higher than 91 % of clips in this corpus. That is a total rise of 0.56.

Nothing was asked of the other axis, and in fact Doubt drifts down from 0.94 to 0.37 (-0.57), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are +0.24, then +0.12, then +0.17, then +0.03 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 57 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.798 before conversion and 0.790 after — it fell by 0.008. Neighbour-to-neighbour the worst pair went 0.584 → 0.645. (The earlier render, with segment 1 left raw, scores 0.742 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Interest moved +0.560 in the original and +0.406 after conversion — 72 % of the delta retained, which is most of it. On the other named axis, Doubt, -0.715 became -0.575.

Quality. Mean predicted overall quality across the segments went 2.81 → 3.05 (+0.24) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.798 → 0.790 -0.008identity cos neighbours 0.584 → 0.645d_b rescored +0.560 → +0.406d_a rescored -0.715 → -0.575d_a mined -0.569d_b mined 0.560min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_W_g-L1gOWlItotal 55.9schain gain +1.5 dBseam step 1.0 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 4 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, slightly relaxed, light breath
(doubt · slow, subdued, steady, monologue) (low mumble) Uhm, so we still have an infrastructure in Berlin about (low mumble) the aspect of (low mumble)
full caption & clip details
An adult masculine voice; delivery is subdued, slow, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as doubt; style: monologue, didactic; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 0.1/10; 6.5s, EN.
EN_W_g-L1gOWlI_W000057 · in -19.3 dBFS · gain -0.7 dB · emolia-02372
(concentration · measured, subdued, steady, monologue) (low mumble) Uh, ecological construction, so, uh, (low mumble) we have a small department in our, uh, (low mumble) city administration in our Berlin Senate of Ecological Construction, and they finance the evaluation of innovative projects.
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: monologue; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 0.6/10; 13.8s, EN.
EN_W_g-L1gOWlI_W000058 · in -20.4 dBFS · gain +0.4 dB · emolia-02372
(normal-paced, normally alert, fairly steady, monologue) This is the Institute of Physics of the Humboldt University, so a different university, it's not mine. (low mumble) Uh, there we harvest rainwater from the non-green roofs, (low mumble) uh, and we irrigate 500, uh, (low mumble) climbing plants in front of the building.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 0.9/10; 14.7s, EN.
EN_W_g-L1gOWlI_W000059 · in -19.1 dBFS · gain -0.8 dB · emolia-02372
(concentration · normal-paced, normally alert, steady, monologue) In, (low mumble) uh, they're growing out of 150 planter boxes. For example, we have some planter boxes with terra preta. So we test different types of soil, different types of irrigation, and we're using the rainwater and eight air conditioners for evaporative exhaust air cooling. This is just a view.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: monologue; average recording, quiet background; genuineness 2.4/6; vocal-burst blend 1.5/10; 18.1s, EN.
EN_W_g-L1gOWlI_W000060 · in -20.1 dBFS · gain +0.1 dB · emolia-02372
(interest · normal-paced, normally alert, fairly steady, casual) Quite interesting of the primary energy demand of the building.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as interest; style: casual, formal; good recording, no background noise; genuineness 2.0/6; vocal-burst blend 1.1/10; 3.6s, EN.
EN_W_g-L1gOWlI_W000061 · in -19.0 dBFS · gain -1.0 dB · emolia-02372
Malevolence Malice(unconstrained axis: Intoxication Altered States of Consciousness)identity −0.00 emotion REVERSED   k-B1-k5 · #13

This chain comes from the one-sided rule: only Malevolence Malice had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Malevolence Malice clearly present — 0.67, higher than 67 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.26.

Nothing was asked of the other axis, and in fact Intoxication Altered States of Consciousness barely moves at all, sitting near 0.88 throughout.

It takes 5 clips to get there. Clip to clip the moves are +0.22, then -0.02, then +0.01, then +0.05 — not a clean run: step 2 moves back the other way by 0.02 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 52 s · zh · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.749 before conversion and 0.746 after — it fell by 0.003. Neighbour-to-neighbour the worst pair went 0.828 → 0.815. (The earlier render, with segment 1 left raw, scores 0.680 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

The emotional move did not survive. Re-scored end to end, Malevolence Malice moved +0.264 in the original and -0.559 after conversion — it changed direction. On this chain the corrected audio is not an improvement. On the other named axis, Intoxication Altered States of Consciousness, -0.048 became -0.223.

Quality. Mean predicted overall quality across the segments went 2.75 → 3.29 (+0.54) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.749 → 0.746 -0.003identity cos neighbours 0.828 → 0.815d_b rescored +0.264 → -0.559d_a rescored -0.048 → -0.223d_a mined -0.048d_b mined 0.264min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00070_S08001total 51.1schain gain +3.5 dBseam step 1.8 dBcrossfades 100/100/100/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · balanced body, average recording, measured, normally alert, slightly relaxed
(moderately variable, some disfluency, average clarity, storytelling) 跟现在都着寒星在一个时候啊,会分为三个呢,小进那个动力呢。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: storytelling, casual; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 1.9/10; 6.5s, ZH.
ZH_B00070_S08001_W000008 · in -22.9 dBFS · gain +2.9 dB · emolia-03973
(fairly steady, frequent disfluency, average clarity, didactic) 现在的亲生太赶这个王姑了,我们反复的影响那个群力家庭。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, formal; average recording, quiet background; genuineness 1.8/6; vocal-burst blend 0.0/10; 6.7s, ZH.
ZH_B00070_S08001_W000009 · in -20.6 dBFS · gain +0.6 dB · emolia-03973
(sourness, anger, contempt · fairly steady, frequent disfluency, average clarity, cartoonish) 这怪太夫臣反了,共同呼作去的东西,要犯太缸,这位厨子等等的很搞鬼,插手求证到某种惩罚,那就做傻子的犯人淋近那个街瓦一楼的。
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, rough, balanced body; average clarity, frequent disfluency, wide pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; reads as sourness, anger, contempt; style: cartoonish, monologue; average recording, quiet background; genuineness 0.9/6; vocal-burst blend 0.8/10; 18.0s, ZH.
ZH_B00070_S08001_W000010 · in -24.1 dBFS · gain +4.1 dB · emolia-03973
(fatigue exhaustion · moderately variable, almost no disfluency, clear, cartoonish) 就警干到大雨,雷大雨占了半夜是一平。
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, balanced body; clear, almost no disfluency, wide pitch range, audible breath; affect is neutral, slightly dominant, fairly guarded; reads as fatigue exhaustion; style: cartoonish, authoritative; average recording, no background noise; genuineness 1.3/6; vocal-burst blend 2.2/10; 10.5s, ZH.
ZH_B00070_S08001_W000011 · in -22.1 dBFS · gain +2.1 dB · emolia-03973
(malevolence malice, anger · steady, frequent disfluency, somewhat unclear, didactic) 都是病人心临意域就可以了。用名字好起洛阳地区发生地震,通宵海水泛滥,波。
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as malevolence malice, anger; style: didactic, monologue; average recording, no background noise; genuineness 1.3/6; vocal-burst blend 0.1/10; 10.0s, ZH.
ZH_B00070_S08001_W000012 · in -21.8 dBFS · gain +1.8 dB · emolia-03973
Impatience and Irritability(unconstrained axis: Pride)identity +0.02 emotion 88 %   k-B1-k5 · #14

This chain comes from the one-sided rule: only Impatience and Irritability had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Impatience and Irritability around average — 0.55, higher than 55 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.38.

Nothing was asked of the other axis, and in fact Pride drifts down from 0.99 to 0.02 (-0.97), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are +0.14, then +0.25, then -0.03, then +0.03 — not a clean run: step 3 moves back the other way by 0.03 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the eurospeech clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 80 s · da · eurospeech

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.868 before conversion and 0.888 after — it rose by 0.020. Neighbour-to-neighbour the worst pair went 0.885 → 0.906. (The earlier render, with segment 1 left raw, scores 0.780 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Impatience and Irritability moved +0.385 in the original and +0.338 after conversion — 88 % of the delta retained, which is most of it. On the other named axis, Pride, -0.970 became -0.578.

Quality. Mean predicted overall quality across the segments went 3.02 → 3.35 (+0.33) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.868 → 0.888 +0.020identity cos neighbours 0.885 → 0.906d_b rescored +0.385 → +0.338d_a rescored -0.970 → -0.578d_a mined -0.970d_b mined 0.385min_cos_consec (site) —min_cos_anchor (site) —dataset eurospeechlang daspeaker denmark_20161M024_2016-11-total 78.2schain gain +1.2 dBseam step 3.2 dBcrossfades 150/150/150/100 ms
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: a middle-aged feminine voice · fairly smooth, average recording, moderately variable, wide pitch range
(pride, triumph, hope enthusiasm optimism · measured, energised, neutral tension, dramatic) Jeg synes jo, det er en rigtig dejlig dag i dag, fordi hele folketinget, alle Folketingets partier, bakker op om en udbygning af den grønne omstilling i Danmark.
full caption & clip details
A middle-aged feminine voice; delivery is energised, measured, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; average clarity, frequent disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as pride, triumph, hope enthusiasm optimism; style: dramatic, cartoonish; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 1.3/10; 14.0s, DA.
denmark_20161M024_2016-11-29_1300_9401345_9415360 · in -24.1 dBFS · gain +4.1 dB · eurospeech-00277
(intoxication altered states of consciousness, disgust, shame · normal-paced, energised, neutral tension, playful) Der er blevet sagt mange grumme ord om vindmøller og havvindmølleparker og kystnære møller i det sidste halve år, (low mumble) og jeg har spændt fulgt debatten og har håbet, at det ikke betød, at vi slækkede på ambitionerne. Jeg skal lade være med at citere diverse
full caption & clip details
A child feminine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as intoxication altered states of consciousness, disgust, shame; style: playful, dramatic; average recording, some background noise; genuineness 3.8/6; vocal-burst blend 3.2/10; 16.7s, DA.
denmark_20161M024_2016-11-29_1300_9415360_9432096 · in -21.4 dBFS · gain +1.4 dB · eurospeech-00277
(jealousy and envy, bitterness, sourness · brisk, energised, neutral tension, cartoonish) folk for, hvad de har sagt om de kystnære møller ved Vesterhav Syd og Vesterhav Nord, men blot glæde mig over, at Venstres borgmester i Ringkøbing-Skjern nu får ret. Han får (ahem) muligheden for at etablere flere tusinde arbejdspladser i det vestjyske, ved at man netop siger ja til de kystnære møller.
full caption & clip details
A child feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, fairly guarded; reads as jealousy and envy, bitterness, sourness; style: cartoonish, ranting; average recording, some background noise; genuineness 2.9/6; vocal-burst blend 2.5/10; 19.0s, DA.
denmark_20161M024_2016-11-29_1300_9432096_9451088 · in -22.8 dBFS · gain +2.8 dB · eurospeech-00277
(disappointment, shame, sadness · brisk, energised, neutral tension, dramatic) Der er ingen tvivl om, at når man laver store beslutninger – og det er både de kystnære møller ved Vesterhav Syd og Vesterhav Nord, det er Kriegers Flak, og det er mange af de grønne investeringer, vi laver – har det selvfølgelig vidtrækkende konsekvenser, også for de mennesker, der bor
full caption & clip details
A child feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as disappointment, shame, sadness; style: dramatic, monologue; average recording, some background noise; genuineness 2.9/6; vocal-burst blend 2.0/10; 16.8s, DA.
denmark_20161M024_2016-11-29_1300_9451088_9467856 · in -23.2 dBFS · gain +3.2 dB · eurospeech-00277
(impatience and irritability, confusion · brisk, normally alert, slightly relaxed, didactic) bor i nærheden af det eller skal se på det, og det er derfor, det for SF har været så afgørende, at man fastholdt den grønne ordning, og at den også skal bruges i forbindelse med de kystnære møller, så netop dem, der bliver påvirket af det,
full caption & clip details
An adult feminine voice; delivery is normally alert, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as impatience and irritability, confusion; style: didactic, ranting; average recording, quiet background; genuineness 2.1/6; vocal-burst blend 0.8/10; 12.4s, DA.
denmark_20161M024_2016-11-29_1300_9467856_9480272 · in -22.9 dBFS · gain +2.9 dB · eurospeech-00277
Concentration(unconstrained axis: Relief)identity −0.02 emotion 109 %   k-B1-k5 · #15

This chain comes from the one-sided rule: only Concentration had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Concentration around average — 0.49, right about the corpus median — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.45.

Nothing was asked of the other axis, and in fact Relief drifts down from 0.83 to 0.45 (-0.38), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are +0.22, then +0.19, then -0.16, then +0.19 — not a clean run: step 3 moves back the other way by 0.16 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 68 s · zh · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.812 before conversion and 0.793 after — it fell by 0.019. Neighbour-to-neighbour the worst pair went 0.882 → 0.868. (The earlier render, with segment 1 left raw, scores 0.726 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.445 in the original and +0.488 after conversion — 109 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Relief, -0.382 became -0.015.

Quality. Mean predicted overall quality across the segments went 3.00 → 3.24 (+0.24) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.812 → 0.793 -0.019identity cos neighbours 0.882 → 0.868d_b rescored +0.445 → +0.488d_a rescored -0.382 → -0.015d_a mined -0.382d_b mined 0.446min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00009_S09777total 67.0schain gain +3.5 dBseam step 2.4 dBcrossfades 100/150/100/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, normally alert, slightly relaxed, fairly steady
(fast, average clarity, casual, playful) 难与不难,并不是衡量风水大师厉害的唯一标准,懂我意思吧?
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, playful; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 4.0/10; 4.9s, ZH.
ZH_B00009_S09777_W000019 · in -21.5 dBFS · gain +1.5 dB · emolia-03369
(doubt · measured, average clarity, authoritative, didactic) 但是你不拿你可以做一个什么无意义的竞争区隔出来。有大风水师很厉害,他也拿罗盘。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as doubt; style: authoritative, didactic; average recording, no background noise; genuineness 2.8/6; vocal-burst blend 3.7/10; 8.4s, ZH.
ZH_B00009_S09777_W000020 · in -20.7 dBFS · gain +0.7 dB · emolia-03369
(doubt, contemplation, jealousy and envy · fast, average clarity, monologue, authoritative) 所以这是人为包装出来的这就是营销的功夫了。他故意不拿呀,人家认为他很厉害,别人都不拿,别人都拿,就他不拿,他一定很厉害,对不对?这就是去安利一种这样子的一种给别人的一个概念。后来他的知名度呢就很高,到最后呢他叫啥名,没人记得他。
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as doubt, contemplation, jealousy and envy; style: monologue, authoritative; average recording, quiet background; genuineness 2.8/6; vocal-burst blend 7.9/10; 19.3s, ZH.
ZH_B00009_S09777_W000021 · in -19.9 dBFS · gain -0.1 dB · emolia-03369
(jealousy and envy, relief, contemplation · normal-paced, average clarity, monologue, authoritative) 别人一见他面就说,哎呀,活罗盘来了,活罗盘来了,他在行业里头就挣了好多的钱。各位觉得这个案例怎么样,有没有启发?有没有启发?所以说定位它重不重要,太重要了。一个好的定位让你赚大钱呢,我在这方面呢印象非常深刻,但是定位是不是独立存在的?不是的啊,定位它是跟命名放在一起的。
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as jealousy and envy, relief, contemplation; style: monologue, authoritative; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 7.4/10; 24.7s, ZH.
ZH_B00009_S09777_W000022 · in -20.4 dBFS · gain +0.3 dB · emolia-03369
(concentration · measured, somewhat unclear, didactic, monologue) 让别人传播,让别人印象深刻也是白搭嘛。所以各位记住喽,命名其实就是定位的一种,也是最简单的定位。好。
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: didactic, monologue; average recording, no background noise; genuineness 3.1/6; vocal-burst blend 3.8/10; 10.4s, ZH.
ZH_B00009_S09777_W000023 · in -21.8 dBFS · gain +1.8 dB · emolia-03369
Concentration(unconstrained axis: Disgust)identity +0.11 emotion 119 %   k-B1-k5 · #16

This chain comes from the one-sided rule: only Concentration had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Concentration around average — 0.45, lower than 55 % of clips in this corpus — and ends with it strongly present at 0.79, higher than 79 % of clips in this corpus. That is a total rise of 0.34.

Nothing was asked of the other axis, and in fact Disgust drifts down from 0.89 to 0.33 (-0.56), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are +0.08, then +0.10, then +0.15, then +0.01 — a plateau around step 4, where it barely moves.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.80 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.80 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 39 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.566 before conversion and 0.680 after — it rose by 0.114. Neighbour-to-neighbour the worst pair went 0.566 → 0.665. (The earlier render, with segment 1 left raw, scores 0.574 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.338 in the original and +0.401 after conversion — 119 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Disgust, -0.559 became -0.542.

Quality. Mean predicted overall quality across the segments went 2.72 → 2.83 (+0.11) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.566 → 0.680 +0.114identity cos neighbours 0.566 → 0.665d_b rescored +0.338 → +0.401d_a rescored -0.559 → -0.542d_a mined -0.559d_b mined 0.338min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_B00017_S09791total 37.6schain gain +1.4 dBseam step 2.2 dBcrossfades 100/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · good recording, no background noise, clear
(slow, very low-energy, relaxed, whispered) We have the longer EE vowel sound in prestige.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, slow, relaxed, steady; timbre is slightly warm, neutral-bright, smooth, thin; clear, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly positive, slightly submissive, neutral openness; no dominant emotion; style: whispered, monologue; good recording, no background noise; genuineness 1.8/6; vocal-burst blend 0.7/10; 6.8s, EN.
EN_B00017_S09791_W000027 · in -25.8 dBFS · gain +5.8 dB · emolia-00575
(awe, contempt, disgust · measured, very low-energy, slightly relaxed, whispered) Not mischievous. There is only one eye here. It's mischievous.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, measured, slightly relaxed, steady; timbre is slightly cool, neutral-bright, fairly smooth, thin; clear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as awe, contempt, disgust; style: whispered, monologue; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 0.0/10; 7.1s, EN.
EN_B00017_S09791_W000028 · in -24.0 dBFS · gain +4.0 dB · emolia-00575
(normal-paced, normally alert, slightly relaxed, whispered) And it's often spelt incorrectly too, pronounced and spelt incorrectly.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: whispered, monologue; good recording, no background noise; genuineness 1.2/6; vocal-burst blend 0.7/10; 4.9s, EN.
EN_B00017_S09791_W000029 · in -28.9 dBFS · gain +8.8 dB · emolia-00575
(malevolence malice, infatuation, contempt · normal-paced, normally alert, slightly relaxed, casual) So this is an adjective that describes a person, usually a child, who is having fun by causing trouble. They're cheeky, it's kind of silly behaviour, not really a negative thing.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, some disfluency, wide pitch range, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as malevolence malice, infatuation, contempt; style: casual, playful; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 0.2/10; 14.6s, EN.
EN_B00017_S09791_W000030 · in -22.7 dBFS · gain +2.7 dB · emolia-00575
(normal-paced, normally alert, slightly relaxed, monologue) Now mischief is behaviour that causes trouble or disruption.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: monologue, casual; good recording, no background noise; genuineness 1.4/6; vocal-burst blend 0.5/10; 5.0s, EN.
EN_B00017_S09791_W000031 · in -21.3 dBFS · gain +1.3 dB · emolia-00575
Astonishment Surprise(unconstrained axis: Affection)identity +0.21 emotion 142 %   k-B1-k5 · #17

This chain comes from the one-sided rule: only Astonishment Surprise had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Astonishment Surprise clearly present — 0.66, higher than 66 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.32.

Nothing was asked of the other axis, and in fact Affection barely moves at all, sitting near 0.98 throughout.

It takes 5 clips to get there. Clip to clip the moves are +0.00, then +0.19, then +0.00, then +0.13 — a plateau around step 1, where it barely moves.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 44 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.321 before conversion and 0.530 after — it rose by 0.209. Neighbour-to-neighbour the worst pair went 0.386 → 0.620. (The earlier render, with segment 1 left raw, scores 0.509 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Astonishment Surprise moved +0.323 in the original and +0.461 after conversion — 142 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Affection, -0.047 became -0.059.

Quality. Mean predicted overall quality across the segments went 2.93 → 2.99 (+0.06) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.321 → 0.530 +0.209identity cos neighbours 0.386 → 0.620d_b rescored +0.323 → +0.461d_a rescored -0.047 → -0.059d_a mined -0.047d_b mined 0.323min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_B00056_S01568total 42.6schain gain +1.5 dBseam step 1.0 dBcrossfades 150/150/150/100 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a child feminine voice · neutral-toned, fairly smooth, good recording, moderately variable, wide pitch range, light breath
(affection, hope enthusiasm optimism · brisk, energised, neutral tension, storytelling) He was a dreamer, and he wanted to make a lot of money.
full caption & clip details
A child feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, no disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as affection, hope enthusiasm optimism; style: storytelling, dramatic; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 7.7/10; 3.7s, EN.
EN_B00056_S01568_W000009 · in -21.6 dBFS · gain +1.6 dB · emolia-01330
(sadness, relief, distress · measured, normally alert, slightly relaxed, storytelling) He remained unemployed and stuck at home. So mom had to provide for the whole family on her own.
full caption & clip details
A child feminine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, neutral openness; reads as sadness, relief, distress; style: storytelling, narration; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 1.4/10; 6.5s, EN.
EN_B00056_S01568_W000010 · in -22.9 dBFS · gain +2.9 dB · emolia-01330
(fatigue exhaustion, longing, sadness · brisk, energised, neutral tension, storytelling) She would come home tired after a hard day's work. And my father would start talking to her about another business venture. And he would tell her that we were all going to be rich soon.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, almost no disfluency, wide pitch range, light breath; affect is negative, slightly dominant, neutral openness; reads as fatigue exhaustion, longing, sadness; style: storytelling, dramatic; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 3.6/10; 9.8s, EN.
EN_B00056_S01568_W000011 · in -20.3 dBFS · gain +0.3 dB · emolia-01330
(distress, helplessness, sadness · normal-paced, normally alert, slightly relaxed, storytelling) Do I even have to tell you that that's not what my mom wanted to hear? Of course they fought again.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as distress, helplessness, sadness; style: storytelling, whispered; good recording, quiet background; genuineness 1.6/6; vocal-burst blend 0.0/10; 6.4s, EN.
EN_B00056_S01568_W000012 · in -24.5 dBFS · gain +4.5 dB · emolia-01330
(astonishment surprise, infatuation, contentment · brisk, energised, neutral tension, playful) And at some point, I have even gotten used to it. But the worst part was when Dad's enthusiasm went beyond his usual tales of wealth, and he really started to create something. First, he created a website with legal advice, which didn't work for even a month.
full caption & clip details
A child feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, little disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as astonishment surprise, infatuation, contentment; style: playful, conversational; good recording, no background noise; genuineness 1.4/6; vocal-burst blend 3.4/10; 17.0s, EN.
EN_B00056_S01568_W000013 · in -21.7 dBFS · gain +1.7 dB · emolia-01330
Doubt(unconstrained axis: Contempt)identity +0.36 emotion 67 %   k-B1-k5 · #18

This chain comes from the one-sided rule: only Doubt had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Doubt clearly present — 0.64, higher than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.32.

Nothing was asked of the other axis, and in fact Contempt drifts down from 0.99 to 0.85 (-0.15), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are +0.24, then +0.05, then +0.07, then -0.03 — not a clean run: step 4 moves back the other way by 0.03 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.12 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.06 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.12, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 57 s · en · podcast

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.235 before conversion and 0.595 after — it rose by 0.360. Neighbour-to-neighbour the worst pair went 0.212 → 0.569. (The earlier render, with segment 1 left raw, scores 0.478 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Doubt moved +0.323 in the original and +0.217 after conversion — 67 % of the delta retained. On the other named axis, Contempt, -0.147 became -0.114.

Quality. Mean predicted overall quality across the segments went 2.72 → 3.02 (+0.30) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.235 → 0.595 +0.360identity cos neighbours 0.212 → 0.569d_b rescored +0.323 → +0.217d_a rescored -0.147 → -0.114d_a mined -0.146d_b mined 0.323min_cos_consec (site) 0.0628min_cos_anchor (site) 0.1231dataset podcastlang enspeaker 679484total 55.5schain gain +3.3 dBseam step 1.5 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · average recording, wide pitch range
(contempt, disgust, malevolence malice · slow, energised, neutral tension, casual) shouldn't get exactly what she asked for. Prosperity Kincaid very wrongfully assumed that Rio and Ali had more control over Pluto than they actually do.
full caption & clip details
An adult masculine voice; delivery is energised, slow, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, balanced body; average clarity, frequent disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as contempt, disgust, malevolence malice; style: casual, playful; average recording, some background noise; mildly explicit content; genuineness 3.9/6; vocal-burst blend 0.0/10; 14.8s, EN.
679484_00219672 · in -14.6 dBFS · gain -5.4 dB · podcast-03531
(impatience and irritability, anger, amusement · brisk, energised, slightly tense, casual) she thought this is on Prosperity. She should remember that while those two were talking, Pluto is the one who walked into an active hostage situation and said, shut the fuck up.
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, slightly tense, moderately variable; timbre is slightly cool, slightly bright, slightly rough, thin; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, slightly guarded; reads as impatience and irritability, anger, amusement; style: casual, playful; average recording, quiet background; mildly explicit content; genuineness 3.6/6; vocal-burst blend 0.6/10; 11.8s, EN.
679484_00221496 · in -18.1 dBFS · gain -1.9 dB · podcast-03525
(confusion, amusement, helplessness · slow, energised, neutral tension, casual) She thought maybe now that this has gone into a less hostile situation, that maybe the two of you could were were no longer distracted and could have maybe reigned the this this uh goblin in. Yeah.
full caption & clip details
An adult masculine voice; delivery is energised, slow, neutral tension, volatile; timbre is slightly cool, neutral-bright, slightly rough, balanced body; average clarity, frequent disfluency, wide pitch range, normal breath; affect is negative, slightly dominant, vulnerable; reads as confusion, amusement, helplessness; style: casual, storytelling; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 0.0/10; 19.5s, EN.
679484_00223040 · in -16.1 dBFS · gain -3.9 dB · podcast-03536
(doubt, astonishment surprise, helplessness · slow, very low-energy, relaxed, storytelling) how do you solve a problem like Maria?
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, slow, relaxed, moderately variable; timbre is slightly warm, dark, slightly rough, balanced body; slurred, frequent disfluency, wide pitch range, light breath; affect is mildly negative, slightly submissive, neutral openness; reads as doubt, astonishment surprise, helplessness; style: storytelling, conversational; average recording, no background noise; genuineness 1.9/6; vocal-burst blend 2.9/10; 4.3s, EN.
679484_00225472 · in -20.8 dBFS · gain +0.8 dB · podcast-03513
(doubt, shame, embarrassment · normal-paced, normally alert, neutral tension, casual) (ahem) I mean the answer to that is (ahem) uh you pose you you position yourself as the last lawyer alive.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as doubt, shame, embarrassment; style: casual, conversational; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 5.8/10; 5.8s, EN.
679484_00226224 · in -21.1 dBFS · gain +1.1 dB · podcast-03510
Concentration(unconstrained axis: Thankfulness Gratitude)identity −0.10 emotion 73 %   k-B1-k5 · #19

This chain comes from the one-sided rule: only Concentration had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Concentration clearly present — 0.64, higher than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.36.

Nothing was asked of the other axis, and in fact Thankfulness Gratitude drifts down from 0.64 to 0.02 (-0.62), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are +0.11, then +0.24, then -0.05, then +0.06 — not a clean run: step 3 moves back the other way by 0.05 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.77 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.77 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 62 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.655 before conversion and 0.554 after — it fell by 0.101. Neighbour-to-neighbour the worst pair went 0.730 → 0.511. (The earlier render, with segment 1 left raw, scores 0.476 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.363 in the original and +0.264 after conversion — 73 % of the delta retained, which is most of it. On the other named axis, Thankfulness Gratitude, -0.623 became -0.606.

Quality. Mean predicted overall quality across the segments went 2.33 → 2.92 (+0.59) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.655 → 0.554 -0.101identity cos neighbours 0.730 → 0.511d_b rescored +0.363 → +0.264d_a rescored -0.623 → -0.606d_a mined -0.623d_b mined 0.363min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_RzTp7te2ORctotal 60.2schain gain +2.3 dBseam step 1.7 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, slightly bright, balanced body, good recording, quiet background, slightly relaxed, some disfluency
(normal-paced, normally alert, fairly steady, casual) Remember that pi is a constant. The derivative of pi is also zero.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; no dominant emotion; style: casual, dramatic; good recording, quiet background; genuineness 2.4/6; vocal-burst blend 1.4/10; 4.2s, EN.
EN_RzTp7te2ORc_W000021 · in -18.2 dBFS · gain -1.8 dB · emolia-02533
(disgust, contempt · normal-paced, energised, moderately variable, casual) I hope you remember it's negative cosine. You can double check that. Cosine by itself.
full caption & clip details
A young adult feminine voice; delivery is energised, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as disgust, contempt; style: casual, storytelling; good recording, quiet background; genuineness 2.8/6; vocal-burst blend 2.3/10; 5.4s, EN.
EN_RzTp7te2ORc_W000023 · in -17.9 dBFS · gain -2.0 dB · emolia-02533
(concentration · brisk, normally alert, fairly steady, casual) So the minus sign in front flips the derivative to be positive. Now how about this plus c? What's a function that after I take its derivative I get just a constant? The answer is plus cx. You can double check that by taking the derivative. Say if I had plus 2x. The derivative is 2 plus 3x.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, minimal breath; affect is positive, slightly dominant, slightly guarded; reads as concentration; style: casual, dramatic; good recording, quiet background; genuineness 0.0/6; vocal-burst blend 1.2/10; 20.6s, EN.
EN_RzTp7te2ORc_W000024 · in -18.8 dBFS · gain -1.2 dB · emolia-02533
(concentration, contentment · brisk, normally alert, fairly steady, casual) Well that's another constant. Usually we will use a different letter for the second constant. Now let's double check that this works. The derivative of negative cosine is sine.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as concentration, contentment; style: casual, dramatic; good recording, quiet background; genuineness 0.9/6; vocal-burst blend 1.4/10; 11.0s, EN.
EN_RzTp7te2ORc_W000026 · in -17.9 dBFS · gain -2.1 dB · emolia-02533
(concentration, interest · normal-paced, normally alert, fairly steady, casual) The derivative of cx is c. The derivative of d is zero. So it works. Normally we use these curly symbols for antiderivatives. So let's translate our chains into curly symbols. So the curly symbol you can read as an English word, the antiderivative of. Here we have the antiderivative of.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, normal breath; affect is mildly positive, slightly dominant, neutral openness; reads as concentration, interest; style: casual, dramatic; good recording, quiet background; genuineness 0.9/6; vocal-burst blend 0.8/10; 19.8s, EN.
EN_RzTp7te2ORc_W000027 · in -18.4 dBFS · gain -1.6 dB · emolia-02533
Interest(unconstrained axis: Disgust)identity +0.34 emotion 115 %   k-B1-k5 · #20

This chain comes from the one-sided rule: only Interest had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Interest clearly present — 0.59, higher than 59 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.37.

Nothing was asked of the other axis, and in fact Disgust drifts down from 0.99 to 0.53 (-0.46), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are +0.03, then +0.03, then +0.07, then +0.24 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.40 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.31 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.40, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 71 s · en · podcast

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.449 before conversion and 0.784 after — it rose by 0.335. Neighbour-to-neighbour the worst pair went 0.298 → 0.850. (The earlier render, with segment 1 left raw, scores 0.677 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Interest moved +0.366 in the original and +0.421 after conversion — 115 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Disgust, -0.457 became -0.251.

Quality. Mean predicted overall quality across the segments went 3.10 → 3.19 (+0.09) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.449 → 0.784 +0.335identity cos neighbours 0.298 → 0.850d_b rescored +0.366 → +0.421d_a rescored -0.457 → -0.251d_a mined -0.457d_b mined 0.369min_cos_consec (site) 0.3128min_cos_anchor (site) 0.4030dataset podcastlang enspeaker 566117total 69.6schain gain +4.1 dBseam step 1.8 dBcrossfades 100/150/150/150 ms
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, no background noise
(disgust, fear · normal-paced, normally alert, slightly relaxed, casual) And at the back door, when you're leaving or entering the house, there's like there's that bottom of it. It's called a
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, slightly dominant, neutral openness; reads as disgust, fear; style: casual, conversational; good recording, no background noise; genuineness 2.9/6; vocal-burst blend 2.9/10; 7.6s, EN.
566117_00131952 · in -25.2 dBFS · gain +5.2 dB · podcast-03217
(confusion, impatience and irritability, anger · normal-paced, normally alert, neutral tension, casual) It's called a threshold. You got you got to cross something. Like you can't, you can't be in the house and out of the house at the same time. It's it's it's almost like it's a necessary shift from one to the other.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, normal breath; affect is mildly negative, neutral stance, neutral openness; reads as confusion, impatience and irritability, anger; style: casual, conversational; average recording, no background noise; mildly explicit content; genuineness 4.9/6; vocal-burst blend 5.9/10; 12.4s, EN.
566117_00132776 · in -23.2 dBFS · gain +3.2 dB · podcast-03195
(fatigue exhaustion, anger · slow, very low-energy, neutral tension, casual) And I think the whole point, previous episode, we were talking about, hey, look, you it's best time to move forward is now
full caption & clip details
An adult masculine voice; delivery is very low-energy, slow, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is mildly negative, neutral stance, slightly guarded; reads as fatigue exhaustion, anger; style: casual, monologue; good recording, no background noise; genuineness 2.6/6; vocal-burst blend 1.4/10; 10.5s, EN.
566117_00134016 · in -23.9 dBFS · gain +3.9 dB · podcast-03221
(contemplation, awe, contentment · slow, normally alert, slightly relaxed, monologue) sometimes you you gotta make that shift. And it's easy to just get kind of stuck where you are at any season of life. And sometimes what should be a season of life, it kind of becomes life, if that makes sense. A season in the wilderness. Because I mean you look at it, the wilderness wasn't so bad.
full caption & clip details
An adult masculine voice; delivery is normally alert, slow, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, normal breath; affect is mildly negative, neutral stance, slightly guarded; reads as contemplation, awe, contentment; style: monologue, casual; good recording, no background noise; genuineness 2.7/6; vocal-burst blend 4.6/10; 22.5s, EN.
566117_00135068 · in -24.2 dBFS · gain +4.2 dB · podcast-06184
(interest, pain, hope enthusiasm optimism · normal-paced, normally alert, neutral tension, casual) Well, and so here's here's kind of what you've put into my my head with this is you know, you hear the good old days. (ahem) Uh people talk about the good old days, and we get caught in the middle, like what you're saying, in the wilderness. And the wilderness may be giving us enough,
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, slightly guarded; reads as interest, pain, hope enthusiasm optimism; style: casual, conversational; good recording, no background noise; genuineness 4.1/6; vocal-burst blend 5.1/10; 17.2s, EN.
566117_00137512 · in -25.0 dBFS · gain +5.0 dB · podcast-03195