emotion__B1__T0.80__C0.25__INTERNAL — voice-corrected

Manifest tier. emotion, rule B1, T=0.8, step cap 0.25. Population 2,257 chains (30 h) over 6 corpora. The SHAREABLE variant of this tier (podcast and evasnippets excluded) holds 1,798.

This is the voice-conversion-corrected version of this tier. The original, uncorrected page is still there and unchanged: t_emotion__B1__T0.80__C0.25__INTERNAL.html. Every card below carries both renders so you can switch between them without leaving the page. What was done, and what it measured →
20chains converted
80segments re-voiced
0.664 → 0.704median worst-to-anchor identity cosine
79 %median emotion delta retained
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken.

(underlined, plain · delivery, style) — the tag before the words. Emotions first, then how it is delivered. Underlined descriptors are the ones that change across this chain — anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words are a real non-speech sound, printed where it happens.

The full generated caption for any clip is under “full caption & clip details”. Its perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.

The tags and captions describe the ORIGINAL clips — they are the corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the converted audio are the ones printed in each card's own paragraph and chips.
Contempt(unconstrained axis: Intoxication Altered States of Consciousness)identity +0.01 emotion 44 %   emotion__B1__T0.80__C0.25__INTERNAL · #1

This chain comes from the one-sided rule: only Contempt had to get where it was going, by at least 0.80. The other emotion was left completely free.

The chain starts with Contempt barely there — 0.14, lower than 86 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.84.

Nothing was asked of the other axis, and in fact Intoxication Altered States of Consciousness drifts down from 0.89 to 0.36 (-0.53), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are +0.16, then +0.24, then +0.22, then +0.22 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the eurospeech clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 81 s · da · eurospeech

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.888 before conversion and 0.903 after — it rose by 0.015. Neighbour-to-neighbour the worst pair went 0.925 → 0.916. (The earlier render, with segment 1 left raw, scores 0.819 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contempt moved +0.840 in the original and +0.366 after conversion — 44 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Intoxication Altered States of Consciousness, -0.528 became -0.479.

Quality. Mean predicted overall quality across the segments went 3.21 → 3.45 (+0.24) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.888 → 0.903 +0.015identity cos neighbours 0.925 → 0.916d_b rescored +0.840 → +0.366d_a rescored -0.528 → -0.479d_a mined -0.528d_b mined 0.840min_cos_consec (site) —min_cos_anchor (site) —dataset eurospeechlang daspeaker denmark_20201M093_2021-04-total 79.5schain gain +2.1 dBseam step 1.0 dBcrossfades 100/150/150/150 ms
Script — 5 chunks, 5 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, average recording, quiet background, normally alert, slightly relaxed, fairly steady
(normal-paced, some disfluency, average clarity, conversational) Tak for svaret. (low mumble) Vi vil selvfølgelig meget gerne diskutere detaljerne med (ahem) ressortministeren. Når jeg alligevel har bedt om, at det er udenrigsministeren, der står her i dag, er det jo, fordi det er Udenrigsministeriet, der står for
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: conversational, monologue; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 1.5/10; 11.2s, DA.
denmark_20201M093_2021-04-14_1300_634224_645392 · in -20.1 dBFS · gain +0.1 dB · eurospeech-00425
(intoxication altered states of consciousness, concentration · measured, frequent disfluency, somewhat unclear, formal) alt med traktaten og de overdragne beføjelser osv. Og i mit parti er (ahem) vi meget bekymrede for, at EU-retten ligesom æder sig ind på stadig flere (low mumble) områder, som vi jo unægtelig (low mumble) har set det, siden vores medlemskab (low mumble) begyndte.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as intoxication altered states of consciousness, concentration; style: formal; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 0.1/10; 17.4s, DA.
denmark_20201M093_2021-04-14_1300_645392_662800 · in -21.2 dBFS · gain +1.2 dB · eurospeech-00425
(concentration, disappointment, contemplation · normal-paced, some disfluency, average clarity, monologue) Nu er vi så kommet hertil, hvor det handler (low mumble) om forhold for børnene, og jeg er ikke i tvivl om, at der er mange børn rundtomkring – sikkert især i de tidligere Sovjetlande – som har det skidt og i hvert fald betragtelig værre, end man har (low mumble) i Danmark. Men er det en berettigelse til, at man så pludselig kan begynde at lave garantier?
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, disappointment, contemplation; style: monologue, authoritative; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 0.5/10; 18.1s, DA.
denmark_20201M093_2021-04-14_1300_662800_680944 · in -21.1 dBFS · gain +1.1 dB · eurospeech-00425
(concentration, triumph, interest · normal-paced, some disfluency, average clarity, monologue) (low mumble) Jeg mindes i hvert fald ikke, at det nogen sinde er blevet forelagt for de danske vælgere, (low mumble) at hvis man stemmer ja til en konkret (ahem) traktat, (low mumble) ja, så får (ahem) EU-Kommissionen beføjelse på børneområdet, socialområdet, eller hvad det nu måtte være af den art.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, triumph, interest; style: monologue; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 1.0/10; 18.9s, DA.
denmark_20201M093_2021-04-14_1300_680944_699808 · in -20.8 dBFS · gain +0.8 dB · eurospeech-00425
(contempt, concentration, disgust · measured, frequent disfluency, average clarity, formal) beføjelse, for det er kun et koordinerende tiltag. Og det er jo den nederste (low mumble) grad af de beføjelser, som man har overdraget i traktaten, og jeg er helt opmærksom på, at det så ikke er lovgivning, og at det ikke har en juridisk
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contempt, concentration, disgust; style: formal; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 0.0/10; 14.7s, DA.
denmark_20201M093_2021-04-14_1300_699808_714496 · in -20.6 dBFS · gain +0.6 dB · eurospeech-00425
Fatigue Exhaustion(unconstrained axis: Emotional Numbness)identity −0.10 emotion 103 %   emotion__B1__T0.80__C0.25__INTERNAL · #2

This chain comes from the one-sided rule: only Fatigue Exhaustion had to get where it was going, by at least 0.80. The other emotion was left completely free.

The chain starts with Fatigue Exhaustion essentially absent — 0.07, lower than 93 % of clips in this corpus — and ends with it strongly present at 0.87, higher than 87 % of clips in this corpus. That is a total rise of 0.81.

Nothing was asked of the other axis, and in fact Emotional Numbness drifts down from 0.95 to 0.73 (-0.21), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are +0.23, then +0.23, then +0.25, then +0.11 — an uneven climb, but always in the same direction.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.85 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.85 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 32 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.733 before conversion and 0.631 after — it fell by 0.102. Neighbour-to-neighbour the worst pair went 0.735 → 0.589. (The earlier render, with segment 1 left raw, scores 0.572 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fatigue Exhaustion moved +0.807 in the original and +0.831 after conversion — 103 % of the delta retained, which is essentially all of it. On the other named axis, Emotional Numbness, -0.212 became -0.134.

Quality. Mean predicted overall quality across the segments went 2.75 → 2.90 (+0.15) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.733 → 0.631 -0.102identity cos neighbours 0.735 → 0.589d_b rescored +0.807 → +0.831d_a rescored -0.212 → -0.134d_a mined -0.212d_b mined 0.807min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_bxIsVHtMl1ytotal 30.4schain gain +0.7 dBseam step 2.5 dBcrossfades 150/150/150/100 ms
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: a middle-aged masculine voice · neutral-toned, slightly dark, balanced body, average recording
(emotional numbness, longing · slow, normally alert, slightly relaxed, didactic) The every, everything's this realm until the on finally callback.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, slow, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, longing; style: didactic, monologue; average recording, quiet background; genuineness 1.7/6; vocal-burst blend 0.3/10; 5.9s, EN.
EN_bxIsVHtMl1y_W000294 · in -18.2 dBFS · gain -1.8 dB · emolia-01770
(emotional numbness · measured, normally alert, slightly relaxed, didactic) Uh, (low mumble) and the on-finally callback is that that is just providing a function.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: didactic, formal; average recording, quiet background; genuineness 2.7/6; vocal-burst blend 0.6/10; 4.9s, EN.
EN_bxIsVHtMl1y_W000295 · in -16.0 dBFS · gain -4.0 dB · emolia-01770
(contemplation, doubt, concentration · slow, very low-energy, relaxed, didactic) So the idea that the on finally callback is determining the promise.
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, slow, relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, submissive, neutral openness; reads as contemplation, doubt, concentration; style: didactic, monologue; average recording, quiet background; genuineness 1.7/6; vocal-burst blend 0.4/10; 8.0s, EN.
EN_bxIsVHtMl1y_W000296 · in -15.6 dBFS · gain -4.4 dB · emolia-01770
(fear, emotional numbness · measured, normally alert, slightly relaxed, didactic) Which is apparently the current semantics that they're trying to change. It sounds like the current semantics is insane.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as fear, emotional numbness; style: didactic, whispered; average recording, quiet background; genuineness 2.0/6; vocal-burst blend 0.5/10; 7.5s, EN.
EN_bxIsVHtMl1y_W000297 · in -17.6 dBFS · gain -2.4 dB · emolia-01770
(measured, normally alert, slightly relaxed, casual) And the semantics they're proposing is compellingly the only possible right
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; average recording, no background noise; genuineness 3.3/6; vocal-burst blend 1.7/10; 4.9s, EN.
EN_bxIsVHtMl1y_W000298 · in -17.9 dBFS · gain -2.1 dB · emolia-01770
Intoxication Altered States of Consciousness(unconstrained axis: Pride)identity −0.03 emotion 96 %   emotion__B1__T0.80__C0.25__INTERNAL · #3

This chain comes from the one-sided rule: only Intoxication Altered States of Consciousness had to get where it was going, by at least 0.80. The other emotion was left completely free.

The chain starts with Intoxication Altered States of Consciousness essentially absent — 0.05, lower than 95 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.91.

Nothing was asked of the other axis, and in fact Pride drifts down from 0.99 to 0.76 (-0.23), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are +0.25, then +0.22, then +0.23, then +0.23 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.78 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.78 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 77 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.773 before conversion and 0.741 after — it fell by 0.032. Neighbour-to-neighbour the worst pair went 0.724 → 0.645. (The earlier render, with segment 1 left raw, scores 0.667 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Intoxication Altered States of Consciousness moved +0.915 in the original and +0.874 after conversion — 96 % of the delta retained, which is essentially all of it. On the other named axis, Pride, -0.232 became -0.239.

Quality. Mean predicted overall quality across the segments went 2.77 → 3.06 (+0.30) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.773 → 0.741 -0.032identity cos neighbours 0.724 → 0.645d_b rescored +0.915 → +0.874d_a rescored -0.232 → -0.239d_a mined -0.231d_b mined 0.915min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_6GP50-GLO7ktotal 75.7schain gain +3.0 dBseam step 2.7 dBcrossfades 100/150/150/100 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a middle-aged somewhat feminine voice
(pride, contentment, relief · measured, very low-energy, slightly relaxed, whispered) We're working out our unit quantity of a hundred millilitres. So put in a hundred ml up here, highlighted both of those and drag those down. So let's look at how I calculated the price. We click on this cell at the top.
full caption & clip details
A middle-aged somewhat feminine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is slightly cool, slightly dark, slightly rough, slightly thin; clear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly positive, neutral stance, neutral openness; reads as pride, contentment, relief; style: whispered, didactic; average recording, quiet background; genuineness 1.2/6; vocal-burst blend 1.3/10; 19.4s, EN.
EN_6GP50-GLO7k_W000011 · in -18.6 dBFS · gain -1.4 dB · emolia-02399
(concentration · slow, very low-energy, slightly relaxed, whispered) E3, we can see the formula up here. What I've done is I've used B3, which is the price currently for this particular item, I've then divided that by C3 which is the amount
full caption & clip details
A middle-aged somewhat feminine voice; delivery is very low-energy, slow, slightly relaxed, fairly steady; timbre is slightly cool, slightly dark, slightly rough, slightly thin; clear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly positive, neutral stance, neutral openness; reads as concentration; style: whispered, didactic; average recording, quiet background; genuineness 1.4/6; vocal-burst blend 0.8/10; 14.9s, EN.
EN_6GP50-GLO7k_W000012 · in -23.0 dBFS · gain +3.0 dB · emolia-02399
(slow, very low-energy, slightly relaxed, whispered) That we've got and what this does is this calculates the price per millilitre.
full caption & clip details
A middle-aged somewhat feminine voice; delivery is very low-energy, slow, slightly relaxed, steady; timbre is slightly cool, slightly dark, fairly smooth, thin; average clarity, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly positive, slightly dominant, neutral openness; no dominant emotion; style: whispered, monologue; good recording, no background noise; genuineness 1.7/6; vocal-burst blend 0.2/10; 7.0s, EN.
EN_6GP50-GLO7k_W000013 · in -18.2 dBFS · gain -1.8 dB · emolia-02399
(normal-paced, normally alert, neutral tension, dramatic) So having calculated the price per milliliter, we then multiply that by our new
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; no dominant emotion; style: dramatic, conversational; good recording, quiet background; genuineness 2.1/6; vocal-burst blend 3.7/10; 5.5s, EN.
EN_6GP50-GLO7k_W000014 · in -15.1 dBFS · gain -4.9 dB · emolia-02399
(intoxication altered states of consciousness, fatigue exhaustion · measured, very low-energy, slightly relaxed, ASMR) quantity, which is a hundred millilitres in F3 and that will give us the price per hundred mils. Having put that formula in this top cell here, we can just drag, click on that, click on the corner. Whoops, didn't mean to do that. Just click back. Right, just click on the corner here and drag that down. And what that does is that puts the formula in
full caption & clip details
A young adult somewhat feminine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is slightly cool, slightly dark, slightly rough, slightly thin; average clarity, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as intoxication altered states of consciousness, fatigue exhaustion; style: ASMR, whispered; average recording, quiet background; genuineness 2.3/6; vocal-burst blend 1.7/10; 29.5s, EN.
EN_6GP50-GLO7k_W000015 · in -18.6 dBFS · gain -1.4 dB · emolia-02399
Disgust(unconstrained axis: Intoxication Altered States of Consciousness)identity +0.01 emotion 74 %   emotion__B1__T0.80__C0.25__INTERNAL · #4

This chain comes from the one-sided rule: only Disgust had to get where it was going, by at least 0.80. The other emotion was left completely free.

The chain starts with Disgust barely there — 0.14, lower than 86 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.82.

Nothing was asked of the other axis, and in fact Intoxication Altered States of Consciousness drifts down from 0.66 to 0.58 (-0.08), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are +0.19, then +0.20, then +0.21, then +0.22 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the snippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 23 s · snippets

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.772 before conversion and 0.781 after — it rose by 0.009. Neighbour-to-neighbour the worst pair went 0.717 → 0.659. (The earlier render, with segment 1 left raw, scores 0.585 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Disgust moved +0.822 in the original and +0.609 after conversion — 74 % of the delta retained, which is most of it. On the other named axis, Intoxication Altered States of Consciousness, -0.097 became +0.013.

Quality. Mean predicted overall quality across the segments went 2.69 → 2.71 (+0.02) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.772 → 0.781 +0.009identity cos neighbours 0.717 → 0.659d_b rescored +0.822 → +0.609d_a rescored -0.097 → +0.013d_a mined -0.080d_b mined 0.822min_cos_consec (site) —min_cos_anchor (site) —dataset snippetslang ?speaker batch84_part0_batch84_parttotal 21.9schain gain +1.1 dBseam step 2.0 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed
(normal-paced, fairly steady, no disfluency, narration) the gang has also made close ties with cartels in South America.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, formal; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 1.2/10; 3.6s.
batch84_part0_batch84_part0_chunk_1756_1_1623773 · in -25.7 dBFS · gain +5.7 dB · snippets-01323
(normal-paced, fairly steady, little disfluency, casual) The gang is the second largest mafia in Italy and is certainly amongst the richest organizations in the world.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: casual, monologue; good recording, no background noise; genuineness 2.3/6; vocal-burst blend 2.3/10; 6.2s.
batch84_part0_batch84_part0_chunk_1756_1_1623784 · in -23.8 dBFS · gain +3.8 dB · snippets-01323
(relief, astonishment surprise, thankfulness gratitude · normal-paced, fairly steady, almost no disfluency, monologue) It's been so beneficial that the family rakes in almost five billion dollars a year in revenue.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as relief, astonishment surprise, thankfulness gratitude; style: monologue, storytelling; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 1.5/10; 5.2s.
batch84_part0_batch84_part0_chunk_1756_1_1623880 · in -23.7 dBFS · gain +3.7 dB · snippets-01323
(measured, fairly steady, no disfluency, formal) the naming rights to ten NFL stadiums.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 2.9/10; 3.2s.
batch84_part0_batch84_part0_chunk_1756_1_1623915 · in -24.3 dBFS · gain +4.3 dB · snippets-01323
(disgust, emotional numbness, distress · normal-paced, steady, no disfluency, formal) prostitution, arms trafficking, money laundering and loan sharking.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as disgust, emotional numbness, distress; style: formal, monologue; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 1.0/10; 4.3s.
batch84_part0_batch84_part0_chunk_1756_1_1623960 · in -22.8 dBFS · gain +2.8 dB · snippets-01323
Intoxication Altered States of Consciousness(unconstrained axis: Sourness)identity +0.09 emotion 6 %   emotion__B1__T0.80__C0.25__INTERNAL · #5

This chain comes from the one-sided rule: only Intoxication Altered States of Consciousness had to get where it was going, by at least 0.80. The other emotion was left completely free.

The chain starts with Intoxication Altered States of Consciousness barely there — 0.14, lower than 86 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.85.

Nothing was asked of the other axis, and in fact Sourness drifts down from 1.00 to 0.60 (-0.40), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are +0.24, then +0.15, then +0.24, then +0.22 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.59 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.52 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.59, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 59 s · en · podcast

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.602 before conversion and 0.689 after — it rose by 0.087. Neighbour-to-neighbour the worst pair went 0.515 → 0.587. (The earlier render, with segment 1 left raw, scores 0.529 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Intoxication Altered States of Consciousness moved +0.842 in the original and +0.048 after conversion — 6 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Sourness, -0.401 became -0.428.

Quality. Mean predicted overall quality across the segments went 2.71 → 3.10 (+0.38) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.602 → 0.689 +0.087identity cos neighbours 0.515 → 0.587d_b rescored +0.842 → +0.048d_a rescored -0.401 → -0.428d_a mined -0.401d_b mined 0.851min_cos_consec (site) 0.5239min_cos_anchor (site) 0.5921dataset podcastlang enspeaker 511300total 57.7schain gain +3.0 dBseam step 1.8 dBcrossfades 150/100/100/100 ms
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, light breath
(sourness, bitterness, impatience and irritability · normal-paced, energised, neutral tension, casual) together. No, they really suck. Like (ahem) if name one thing the American military accomplished. They they didn't they didn't stop the Holocaust. The Soviets did that, and that's the biggest thing they take credit for. I mean, sure and then they dropped some fucking nukes on Japan and the people are still suffering
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as sourness, bitterness, impatience and irritability; style: casual, playful; average recording, some background noise; genuineness 6.0/6; vocal-burst blend 7.1/10; 17.2s, EN.
511300_00286720 · in -17.7 dBFS · gain -2.3 dB · podcast-02322
(disgust, anger, contempt · normal-paced, normally alert, neutral tension, conversational) from that shit. And we wanna we wanna go and say that this building that fell over is the biggest act of terror act of terrorism. Like we invented a weapon of mass destruction. We
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, neutral openness; reads as disgust, anger, contempt; style: conversational, casual; average recording, quiet background; genuineness 5.9/6; vocal-burst blend 5.1/10; 9.6s, EN.
511300_00288448 · in -17.9 dBFS · gain -2.1 dB · podcast-02319
(embarrassment · normal-paced, normally alert, relaxed, casual) were the first ones to do it. Technically they were technically they already existed in Germany, but we're talking, you know, the the sarin gas that he had, it wouldn't it couldn't do what that nuke did. And after
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as embarrassment; style: casual, conversational; average recording, quiet background; genuineness 5.2/6; vocal-burst blend 3.3/10; 11.9s, EN.
511300_00289400 · in -19.6 dBFS · gain -0.4 dB · podcast-02323
(amusement · normal-paced, normally alert, slightly relaxed, casual) the after we dropped those nukes, we started limiting things like sarin gas,
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as amusement; style: casual, conversational; average recording, quiet background; genuineness 4.9/6; vocal-burst blend 4.5/10; 3.5s, EN.
511300_00290584 · in -19.8 dBFS · gain -0.2 dB · podcast-02326
(intoxication altered states of consciousness, amusement, embarrassment · measured, normally alert, relaxed, casual) Syria, it happened. (low mumble) Um yeah, I mean (low mumble) the US military. They give some cool planes though. Oh also did you know did you know that the they they suspected that bin Laden was gonna attack on the
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, wide pitch range, light breath; affect is mildly negative, slightly submissive, neutral openness; reads as intoxication altered states of consciousness, amusement, embarrassment; style: casual, playful; average recording, quiet background; genuineness 5.3/6; vocal-burst blend 1.7/10; 16.2s, EN.
511300_00291200 · in -17.6 dBFS · gain -2.4 dB · podcast-02326
Fatigue Exhaustion(unconstrained axis: Concentration)identity +0.45 emotion 97 %   emotion__B1__T0.80__C0.25__INTERNAL · #6

This chain comes from the one-sided rule: only Fatigue Exhaustion had to get where it was going, by at least 0.80. The other emotion was left completely free.

The chain starts with Fatigue Exhaustion barely there — 0.13, lower than 87 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.83.

Nothing was asked of the other axis, and in fact Concentration drifts down from 0.90 to 0.28 (-0.61), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are +0.20, then +0.21, then +0.20, then +0.21 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.09 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.09 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 48 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.185 before conversion and 0.635 after — it rose by 0.450. Neighbour-to-neighbour the worst pair went 0.185 → 0.652. (The earlier render, with segment 1 left raw, scores 0.597 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fatigue Exhaustion moved +0.826 in the original and +0.804 after conversion — 97 % of the delta retained, which is essentially all of it. On the other named axis, Concentration, -0.613 became -0.594.

Quality. Mean predicted overall quality across the segments went 2.68 → 2.97 (+0.29) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.185 → 0.635 +0.450identity cos neighbours 0.185 → 0.652d_b rescored +0.826 → +0.804d_a rescored -0.613 → -0.594d_a mined -0.613d_b mined 0.826min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_t3zsbnPp3Bctotal 46.8schain gain +2.4 dBseam step 1.2 dBcrossfades 150/150/150/100 ms
Script — 5 chunks, 5 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, fairly smooth, balanced body, fairly steady
(normal-paced, normally alert, slightly relaxed, monologue) Is technology creating new spaces for work and changing others that (low mumble) are being (ahem) left redundant from the private sector perspective?
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, didactic; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 0.7/10; 9.4s, EN.
EN_t3zsbnPp3Bc_W000330 · in -22.4 dBFS · gain +2.4 dB · emolia-01571
(slow, subdued, slightly relaxed, monologue) Yeah, (low mumble) the future is globalization. (low mumble) Uh, technology is enabling us, uh, (low mumble) to leverage cross-border trade, uh, (low mumble) drive innovation. (low mumble) Uhm, so it's, it's, (low mumble) uhm,
full caption & clip details
An adult masculine voice; delivery is subdued, slow, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 0.0/10; 11.5s, EN.
EN_t3zsbnPp3Bc_W000331 · in -20.4 dBFS · gain +0.4 dB · emolia-01571
(interest, concentration · normal-paced, normally alert, slightly relaxed, monologue) In some ways it's creating jobs, I know it's taking away jobs, but it's creating different types of jobs. So it reverts back to the big question on education. How do we train the workforce to participate in more of a globalization (ahem) network?
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as interest, concentration; style: monologue, formal; good recording, no background noise; genuineness 1.4/6; vocal-burst blend 1.0/10; 12.4s, EN.
EN_t3zsbnPp3Bc_W000332 · in -20.7 dBFS · gain +0.7 dB · emolia-01571
(intoxication altered states of consciousness, contemplation · measured, subdued, relaxed, casual) My view is, is, is similar. I think, you know, more and more we are a service economy. And, (low mumble) uh, you know, we kind of, you know, say that data is the, is the new wild, so.
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is mildly positive, submissive, neutral openness; reads as intoxication altered states of consciousness, contemplation; style: casual, conversational; average recording, quiet background; genuineness 5.1/6; vocal-burst blend 4.8/10; 10.9s, EN.
EN_t3zsbnPp3Bc_W000333 · in -21.1 dBFS · gain +1.1 dB · emolia-01571
(fatigue exhaustion, intoxication altered states of consciousness · measured, subdued, slightly relaxed, casual) That needs to kind of, you know, really, (low mumble) uh, keep going.
full caption & clip details
A young adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; slurred, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as fatigue exhaustion, intoxication altered states of consciousness; style: casual, monologue; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 1.1/10; 3.5s, EN.
EN_t3zsbnPp3Bc_W000334 · in -18.8 dBFS · gain -1.2 dB · emolia-01571
Contempt(unconstrained axis: Affection)identity +0.14 emotion 68 %   emotion__B1__T0.80__C0.25__INTERNAL · #7

This chain comes from the one-sided rule: only Contempt had to get where it was going, by at least 0.80. The other emotion was left completely free.

The chain starts with Contempt barely there — 0.14, lower than 86 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.81.

Nothing was asked of the other axis, and in fact Affection drifts down from 0.82 to 0.60 (-0.22), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are +0.16, then +0.21, then +0.22, then +0.22 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.07 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.10 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.07, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 43 s · en · podcast

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.146 before conversion and 0.285 after — it rose by 0.138. Neighbour-to-neighbour the worst pair went 0.138 → 0.275. (The earlier render, with segment 1 left raw, scores 0.342 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contempt moved +0.809 in the original and +0.552 after conversion — 68 % of the delta retained. On the other named axis, Affection, -0.218 became +0.000.

Quality. Mean predicted overall quality across the segments went 2.76 → 2.96 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.146 → 0.285 +0.138identity cos neighbours 0.138 → 0.275d_b rescored +0.809 → +0.552d_a rescored -0.218 → +0.000d_a mined -0.222d_b mined 0.808min_cos_consec (site) 0.0988min_cos_anchor (site) 0.0710dataset podcastlang enspeaker 886199total 41.7schain gain +3.8 dBseam step 3.0 dBcrossfades 150/150/150/100 ms
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · fairly smooth, slightly relaxed, fairly steady
(normal-paced, normally alert, frequent disfluency, conversational) (wistful sigh) Uh we have special handling of cats, we use (low mumble) uh local anesthetics to draw the blood from them.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; no dominant emotion; style: conversational, ASMR; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 0.4/10; 6.7s, EN.
886199_00010612 · in -20.5 dBFS · gain +0.5 dB · podcast-05367
(measured, normally alert, no disfluency, monologue) So we are using the most modern ways how to treat cats.
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; no dominant emotion; style: monologue, ASMR; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 5.4/10; 4.2s, EN.
886199_00011756 · in -20.5 dBFS · gain +0.5 dB · podcast-05359
(normal-paced, normally alert, some disfluency, casual) Okay. And do they they come through the same entrance but they just go into a different place?
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, conversational; good recording, no background noise; genuineness 4.6/6; vocal-burst blend 2.1/10; 4.1s, EN.
886199_00012808 · in -20.4 dBFS · gain +0.4 dB · podcast-04643
(normal-paced, normally alert, frequent disfluency, casual) exactly. So are you saying that cats are so sensitive to the smell that if they smell a dog that then they will become much more nervous. So this is the reason why you have that separated area so they feel more comfortable.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, whispered; average recording, quiet background; genuineness 4.8/6; vocal-burst blend 1.2/10; 13.4s, EN.
886199_00013312 · in -20.6 dBFS · gain +0.6 dB · podcast-05360
(contempt, emotional numbness · normal-paced, very low-energy, frequent disfluency, casual) Yeah, yeah. Uh (low mumble) it's not just the smell, it's also vision, see the dog or hear the dog. So they don't see and hear the dogs when they are waiting, because there is completely separated area. Okay,
full caption & clip details
A young adult feminine voice; delivery is very low-energy, normal-paced, slightly relaxed, fairly steady; timbre is slightly cool, slightly bright, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly positive, neutral stance, neutral openness; reads as contempt, emotional numbness; style: casual, monologue; average recording, quiet background; mildly explicit content; genuineness 4.1/6; vocal-burst blend 1.9/10; 14.0s, EN.
886199_00014688 · in -20.7 dBFS · gain +0.7 dB · podcast-05363
Astonishment Surprise(unconstrained axis: Interest)identity +0.52 emotion 90 %   emotion__B1__T0.80__C0.25__INTERNAL · #8

This chain comes from the one-sided rule: only Astonishment Surprise had to get where it was going, by at least 0.80. The other emotion was left completely free.

The chain starts with Astonishment Surprise barely there — 0.14, lower than 86 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.85.

Nothing was asked of the other axis, and in fact Interest drifts down from 0.98 to 0.66 (-0.32), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are +0.22, then +0.16, then +0.22, then +0.24 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.23 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.25 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.23, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 45 s · en · podcast

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.200 before conversion and 0.719 after — it rose by 0.519. Neighbour-to-neighbour the worst pair went 0.259 → 0.836. (The earlier render, with segment 1 left raw, scores 0.661 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Astonishment Surprise moved +0.855 in the original and +0.767 after conversion — 90 % of the delta retained, which is most of it. On the other named axis, Interest, -0.329 became -0.271.

Quality. Mean predicted overall quality across the segments went 2.91 → 3.13 (+0.22) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.200 → 0.719 +0.519identity cos neighbours 0.259 → 0.836d_b rescored +0.855 → +0.767d_a rescored -0.329 → -0.271d_a mined -0.323d_b mined 0.855min_cos_consec (site) 0.2475min_cos_anchor (site) 0.2256dataset podcastlang enspeaker 183770total 43.9schain gain +3.6 dBseam step 1.1 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, good recording, average clarity, light breath
(interest, contemplation, doubt · measured, subdued, relaxed, casual) harnessing that power, trying to capture that lightning in a bottle. And I will say this if they're able to do it with a couple of other shows that are kind of ripe for the picking,
full caption & clip details
An adult masculine voice; delivery is subdued, measured, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as interest, contemplation, doubt; style: casual, conversational; good recording, quiet background; genuineness 4.2/6; vocal-burst blend 4.9/10; 13.5s, EN.
183770_00105768 · in -18.6 dBFS · gain -1.4 dB · podcast-02784
(measured, normally alert, slightly relaxed, whispered) we're gonna be in a good spot. Okay. So I I will say the possibility of us getting a continuation of Spider-Man.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: whispered, conversational; good recording, no background noise; genuineness 2.6/6; vocal-burst blend 2.6/10; 9.7s, EN.
183770_00107352 · in -18.0 dBFS · gain -2.0 dB · podcast-02781
(normal-paced, normally alert, slightly relaxed, casual) Spider-Man 98 has been thrown around. Okay, and we have seen our friend Spider Guy.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; good recording, quiet background; genuineness 3.0/6; vocal-burst blend 3.2/10; 5.1s, EN.
183770_00108480 · in -17.4 dBFS · gain -2.6 dB · podcast-02768
(fatigue exhaustion · normal-paced, normally alert, slightly relaxed, casual) (ahem) um, and in in the the final final season here again, spoiler alert. In in the final episode, rather, we saw Peter Parker and Mary Jane.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as fatigue exhaustion; style: casual, conversational; good recording, quiet background; genuineness 3.5/6; vocal-burst blend 4.2/10; 9.0s, EN.
183770_00109040 · in -17.3 dBFS · gain -2.7 dB · podcast-02813
(astonishment surprise, relief, triumph · normal-paced, normally alert, neutral tension, casual) I did, and then I we watched (ahem) uh a breakdown video, and I was like, Oh, that's why they look familiar. That's okay. Now I got it. There it is.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as astonishment surprise, relief, triumph; style: casual, conversational; good recording, quiet background; genuineness 5.0/6; vocal-burst blend 3.6/10; 7.2s, EN.
183770_00110288 · in -20.1 dBFS · gain +0.1 dB · podcast-02791
Fear(unconstrained axis: Embarrassment)identity −0.02 emotion 98 %   emotion__B1__T0.80__C0.25__INTERNAL · #9

This chain comes from the one-sided rule: only Fear had to get where it was going, by at least 0.80. The other emotion was left completely free.

The chain starts with Fear barely there — 0.17, lower than 83 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.83.

Nothing was asked of the other axis, and in fact Embarrassment barely moves at all, sitting near 0.97 throughout.

It takes 5 clips to get there. Clip to clip the moves are +0.23, then +0.19, then +0.22, then +0.20 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.89 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.89 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 49 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.819 before conversion and 0.803 after — it fell by 0.016. Neighbour-to-neighbour the worst pair went 0.811 → 0.815. (The earlier render, with segment 1 left raw, scores 0.738 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fear moved +0.833 in the original and +0.815 after conversion — 98 % of the delta retained, which is essentially all of it. On the other named axis, Embarrassment, +0.018 became +0.020.

Quality. Mean predicted overall quality across the segments went 2.90 → 3.06 (+0.16) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.819 → 0.803 -0.016identity cos neighbours 0.811 → 0.815d_b rescored +0.833 → +0.815d_a rescored +0.018 → +0.020d_a mined 0.018d_b mined 0.833min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_B00055_S09869total 47.8schain gain +3.7 dBseam step 2.6 dBcrossfades 100/100/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, slightly bright, fairly smooth, moderately variable, some disfluency, wide pitch range
(embarrassment, shame, affection · brisk, normally alert, slightly relaxed, conversational) I'm like, are we good? 100%. I think it's important in any relationship to have arguments and conflict, but I came from a few friendships that...
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as embarrassment, shame, affection; style: conversational, casual; good recording, no background noise; genuineness 3.5/6; vocal-burst blend 7.6/10; 8.2s, EN.
EN_B00055_S09869_W000006 · in -20.8 dBFS · gain +0.8 dB · emolia-01313
(embarrassment, contemplation, longing · brisk, energised, neutral tension, casual) I couldn't argue in them, I couldn't really speak my mind, so I remember when you and I became friends, you were kind of like a bad atta. How? Can we say that?
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as embarrassment, contemplation, longing; style: casual, conversational; good recording, no background noise; genuineness 4.1/6; vocal-burst blend 7.1/10; 8.4s, EN.
EN_B00055_S09869_W000007 · in -20.3 dBFS · gain +0.3 dB · emolia-01313
(embarrassment, pleasure ecstasy, thankfulness gratitude · brisk, energised, neutral tension, casual) You're very vocal in how you're feeling. And so, and I, I told you, I came from these friendships where I felt like I couldn't really speak up. They didn't really, it was just a weird thing where we couldn't really have open communication like that. So when I came into our friendship, do you remember the first
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; somewhat unclear, some disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, slightly vulnerable; reads as embarrassment, pleasure ecstasy, thankfulness gratitude; style: casual, conversational; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 9.6/10; 13.6s, EN.
EN_B00055_S09869_W000008 · in -19.9 dBFS · gain -0.1 dB · emolia-01313
(elation, relief, pleasure ecstasy · brisk, energised, neutral tension, casual) Little thing that we got into the first argument. I thought, I thought, I thought it was over for us. I thought it was way before Girls Gone Bible.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as elation, relief, pleasure ecstasy; style: casual, playful; average recording, quiet background; genuineness 4.8/6; vocal-burst blend 5.7/10; 8.8s, EN.
EN_B00055_S09869_W000009 · in -19.0 dBFS · gain -1.0 dB · emolia-01313
(fear, infatuation, distress · normal-paced, energised, neutral tension, casual) I literally start bawling my eyes out because you said something like, I'm concerned for this friend. You said something like that and it triggered me so hard. I thought you were breaking up with me.
full caption & clip details
A young adult feminine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, slightly vulnerable; reads as fear, infatuation, distress; style: casual, conversational; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 5.9/10; 9.4s, EN.
EN_B00055_S09869_W000010 · in -19.6 dBFS · gain -0.5 dB · emolia-01313
Concentration(unconstrained axis: Hope Enthusiasm Optimism)identity +0.48 emotion 106 %   emotion__B1__T0.80__C0.25__INTERNAL · #10

This chain comes from the one-sided rule: only Concentration had to get where it was going, by at least 0.80. The other emotion was left completely free.

The chain starts with Concentration barely there — 0.15, lower than 85 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.83.

Nothing was asked of the other axis, and in fact Hope Enthusiasm Optimism drifts down from 1.00 to 0.82 (-0.18), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are +0.23, then +0.19, then +0.22, then +0.18 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.12 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.12 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 76 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.263 before conversion and 0.742 after — it rose by 0.479. Neighbour-to-neighbour the worst pair went 0.306 → 0.671. (The earlier render, with segment 1 left raw, scores 0.625 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.828 in the original and +0.875 after conversion — 106 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Hope Enthusiasm Optimism, -0.181 became -0.238.

Quality. Mean predicted overall quality across the segments went 2.87 → 3.15 (+0.28) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.263 → 0.742 +0.479identity cos neighbours 0.306 → 0.671d_b rescored +0.828 → +0.875d_a rescored -0.181 → -0.238d_a mined -0.181d_b mined 0.828min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN__SMs5f7SbzQtotal 74.2schain gain +3.7 dBseam step 3.2 dBcrossfades 150/100/100/150 ms
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, fairly smooth, balanced body, slightly relaxed, fairly steady
(hope enthusiasm optimism, contentment, thankfulness gratitude · brisk, normally alert, some disfluency, casual) Hello and welcome back to a special episode of Community Hotline. My name is James Ofsink and we're here today to talk about the Gresham Safety Measure. (low mumble) Uh, we're fortunate to be joined by Gresham City Manager Nina Vetter and Gresham Fire Chief Scott Lewis. Thank you so much both for being here. Thank you.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, minimal breath; affect is positive, slightly submissive, slightly guarded; reads as hope enthusiasm optimism, contentment, thankfulness gratitude; style: casual, conversational; good recording, quiet background; genuineness 0.9/6; vocal-burst blend 3.4/10; 16.4s, EN.
EN__SMs5f7SbzQ_W000000 · in -15.7 dBFS · gain -4.3 dB · emolia-01709
(doubt · normal-paced, normally alert, some disfluency, casual) Nina, can you tell us a little bit about the primary goals of the upcoming levy?
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as doubt; style: casual, conversational; good recording, quiet background; genuineness 2.8/6; vocal-burst blend 2.9/10; 4.4s, EN.
EN__SMs5f7SbzQ_W000001 · in -13.3 dBFS · gain -6.7 dB · emolia-01709
(hope enthusiasm optimism, contentment, thankfulness gratitude · normal-paced, very low-energy, some disfluency, casual) Yeah, absolutely. So our Gresham safety levy that's on the ballot in May of this year, (low mumble) uhm, is a five-year operating levy that would fund our critical safety services at an average cost of about $28 (ahem) per household. (ahem) Uhm, so our community safety levy is really focused on a few key areas. Police, fire, homelessness response, and mental health response.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, minimal breath; affect is positive, slightly dominant, neutral openness; reads as hope enthusiasm optimism, contentment, thankfulness gratitude; style: casual, monologue; good recording, quiet background; genuineness 0.5/6; vocal-burst blend 1.7/10; 24.7s, EN.
EN__SMs5f7SbzQ_W000002 · in -15.5 dBFS · gain -4.5 dB · emolia-01709
(disappointment, helplessness, sadness · normal-paced, normally alert, some disfluency, monologue) (ahem) Since the city has been in a budget crisis for the last few decades, (ahem) we have not been able to provide the community really with the services (ahem) that they need, that they demand, and the services that are needed in these changing times (ahem) as safety has evolved. So the Gresham Safety Levy would allow us to
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as disappointment, helplessness, sadness; style: monologue, casual; average recording, quiet background; genuineness 1.6/6; vocal-burst blend 1.6/10; 19.0s, EN.
EN__SMs5f7SbzQ_W000003 · in -16.1 dBFS · gain -3.9 dB · emolia-01709
(concentration, pain · brisk, normally alert, almost no disfluency, formal) Stabilize our current services that we do provide in safety and add strategically positions across those critical safety areas so that we can be more responsive and we can be proactive.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as concentration, pain; style: formal, dramatic; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.1/10; 10.4s, EN.
EN__SMs5f7SbzQ_W000004 · in -15.2 dBFS · gain -4.8 dB · emolia-01709
Relief(unconstrained axis: Triumph)identity +0.12 emotion 10 %   emotion__B1__T0.80__C0.25__INTERNAL · #11

This chain comes from the one-sided rule: only Relief had to get where it was going, by at least 0.80. The other emotion was left completely free.

The chain starts with Relief barely there — 0.15, lower than 85 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.82.

Nothing was asked of the other axis, and in fact Triumph drifts down from 0.92 to 0.20 (-0.72), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are +0.21, then +0.23, then +0.18, then +0.21 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.75 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.75 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 61 s · zh · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.689 before conversion and 0.809 after — it rose by 0.120. Neighbour-to-neighbour the worst pair went 0.577 → 0.794. (The earlier render, with segment 1 left raw, scores 0.763 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Relief moved +0.821 in the original and +0.080 after conversion — 10 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Triumph, -0.724 became -0.784.

Quality. Mean predicted overall quality across the segments went 3.01 → 3.22 (+0.21) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.689 → 0.809 +0.120identity cos neighbours 0.577 → 0.794d_b rescored +0.821 → +0.080d_a rescored -0.724 → -0.784d_a mined -0.725d_b mined 0.821min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00047_S01791total 59.8schain gain +1.7 dBseam step 2.0 dBcrossfades 100/150/100/100 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a child feminine voice · average recording, normally alert, moderately variable
(triumph · measured, slightly relaxed, no disfluency, storytelling) 我拿着放着笔的墨水瓶,欢快的哼着歌,一蹦一跳的跑过来。
full caption & clip details
A child feminine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is neutral-toned, bright, fairly smooth, balanced body; average clarity, no disfluency, wide pitch range, light breath; affect is neutral, slightly submissive, neutral openness; reads as triumph; style: storytelling, cartoonish; average recording, no background noise; genuineness 1.9/6; vocal-burst blend 3.0/10; 6.1s, ZH.
ZH_B00047_S01791_W000002 · in -30.1 dBFS · gain +10.1 dB · emolia-03743
(measured, slightly relaxed, some disfluency, cartoonish) 爸爸双手抱住脑袋,沮丧的尖叫着,然后气冲冲的离开了。我知道爸爸去找罐子了怎么办?又要挨打了,我得想办法补救一下。有啦。
full caption & clip details
A child strongly feminine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is slightly cool, bright, fairly smooth, slightly thin; clear, some disfluency, very wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; no dominant emotion; style: cartoonish, storytelling; average recording, no background noise; genuineness 2.5/6; vocal-burst blend 4.7/10; 18.1s, ZH.
ZH_B00047_S01791_W000003 · in -29.7 dBFS · gain +9.7 dB · emolia-03743
(intoxication altered states of consciousness · measured, slightly relaxed, some disfluency, storytelling) 我把墨水瓶放在一边,拿着笔开始在弄脏的地方画画。
full caption & clip details
A child feminine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is neutral-toned, bright, fairly smooth, balanced body; slurred, some disfluency, very wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as intoxication altered states of consciousness; style: storytelling, playful; average recording, no background noise; genuineness 2.8/6; vocal-burst blend 3.4/10; 5.8s, ZH.
ZH_B00047_S01791_W000004 · in -32.3 dBFS · gain +12.3 dB · emolia-03743
(astonishment surprise, doubt, contempt · measured, neutral tension, some disfluency, cartoonish) 臭小子又闯祸,看我怎么收拾你。他举起棍子,正准备打我,却看到了我画的小猴子,忍不住夸道。咦画的不错嘛,爸爸自己也手痒痒了,说看着儿子老爸给你录一首。
full caption & clip details
An elderly feminine voice; delivery is normally alert, measured, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, slightly thin; average clarity, some disfluency, very wide pitch range, normal breath; affect is mildly positive, slightly submissive, neutral openness; reads as astonishment surprise, doubt, contempt; style: cartoonish, storytelling; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 6.9/10; 23.4s, ZH.
ZH_B00047_S01791_W000005 · in -24.8 dBFS · gain +4.8 dB · emolia-03743
(relief · normal-paced, slightly relaxed, some disfluency, storytelling) 说完,他拿起墨水瓶浇洒、墨汁地毯,中间出现了一条蛇。
full caption & clip details
A child feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is cool, bright, fairly smooth, balanced body; crisply articulate, some disfluency, very wide pitch range, light breath; affect is neutral, slightly submissive, neutral openness; reads as relief; style: storytelling, cartoonish; average recording, no background noise; genuineness 1.8/6; vocal-burst blend 2.1/10; 7.2s, ZH.
ZH_B00047_S01791_W000006 · in -31.6 dBFS · gain +11.6 dB · emolia-03743
Triumph(unconstrained axis: Sadness)identity +0.39 emotion 49 %   emotion__B1__T0.80__C0.25__INTERNAL · #12

This chain comes from the one-sided rule: only Triumph had to get where it was going, by at least 0.80. The other emotion was left completely free.

The chain starts with Triumph barely there — 0.08, lower than 92 % of clips in this corpus — and ends with it at the very top of the corpus at 0.91, higher than 91 % of clips in this corpus. That is a total rise of 0.82.

Nothing was asked of the other axis, and in fact Sadness drifts down from 0.98 to 0.43 (-0.55), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are +0.24, then +0.24, then +0.20, then +0.15 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the snippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 29 s · snippets

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.100 before conversion and 0.486 after — it rose by 0.386. Neighbour-to-neighbour the worst pair went -0.010 → 0.505. (The earlier render, with segment 1 left raw, scores 0.449 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Triumph moved +0.824 in the original and +0.400 after conversion — 49 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Sadness, -0.548 became -0.516.

Quality. Mean predicted overall quality across the segments went 2.73 → 2.88 (+0.15) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.100 → 0.486 +0.386identity cos neighbours -0.010 → 0.505d_b rescored +0.824 → +0.400d_a rescored -0.548 → -0.516d_a mined -0.547d_b mined 0.823min_cos_consec (site) —min_cos_anchor (site) —dataset snippetslang ?speaker batch260_part3_batch260_patotal 27.5schain gain +2.0 dBseam step 2.2 dBcrossfades 150/100/100/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, fairly smooth, normally alert, slightly relaxed, moderate pitch range
(sadness, distress, disappointment · brisk, fairly steady, no disfluency, storytelling) This was disappointing as they were putting a lot of effort into finding who had killed Jemima.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as sadness, distress, disappointment; style: storytelling, whispered; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 1.3/10; 5.0s.
batch260_part3_batch260_part3_chunk_808_1_787811 · in -24.5 dBFS · gain +4.5 dB · snippets-00837
(normal-paced, fairly steady, no disfluency, narration) Last month's bonus episode focused on the deep freeze murder of Anne Noblet.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: narration, whispered; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 2.6/10; 4.6s.
batch260_part3_batch260_part3_chunk_808_1_787866 · in -26.2 dBFS · gain +6.2 dB · snippets-00837
(measured, steady, almost no disfluency, narration) they're being used together, for example, using music therapy alongside of chemotherapy for a child with acute lymphocytic leukemia.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, formal; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 0.9/10; 7.0s.
batch260_part3_batch260_part3_chunk_808_1_788156 · in -34.6 dBFS · gain +14.6 dB · snippets-00837
(awe, sourness, emotional numbness · measured, fairly steady, no disfluency, ASMR) you'd say it's a relative who would claim to be a spirit guide or an ascended master.
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as awe, sourness, emotional numbness; style: ASMR, casual; average recording, no background noise; genuineness 2.5/6; vocal-burst blend 2.3/10; 4.7s.
batch260_part3_batch260_part3_chunk_808_1_788226 · in -16.8 dBFS · gain -3.2 dB · snippets-00837
(triumph · fast, fairly steady, some disfluency, casual) you know, but in the case of that boy, I mean we grew up together in a in a in an area in Abeokuta. So, I know him.
full caption & clip details
An adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, thin; slurred, some disfluency, moderate pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; reads as triumph; style: casual, storytelling; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 6.5/10; 6.8s.
batch260_part3_batch260_part3_chunk_808_1_788412 · in -27.7 dBFS · gain +7.7 dB · snippets-00837
Emotional Numbness(unconstrained axis: Elation)identity −0.02 emotion 77 %   emotion__B1__T0.80__C0.25__INTERNAL · #13

This chain comes from the one-sided rule: only Emotional Numbness had to get where it was going, by at least 0.80. The other emotion was left completely free.

The chain starts with Emotional Numbness barely there — 0.13, lower than 87 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.86.

Nothing was asked of the other axis, and in fact Elation drifts down from 0.95 to 0.13 (-0.82), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are +0.20, then +0.22, then +0.21, then +0.23 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.92 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.92 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 54 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.848 before conversion and 0.829 after — it fell by 0.019. Neighbour-to-neighbour the worst pair went 0.857 → 0.815. (The earlier render, with segment 1 left raw, scores 0.743 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.858 in the original and +0.663 after conversion — 77 % of the delta retained, which is most of it. On the other named axis, Elation, -0.824 became -0.071.

Quality. Mean predicted overall quality across the segments went 2.93 → 3.11 (+0.17) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.848 → 0.829 -0.019identity cos neighbours 0.857 → 0.815d_b rescored +0.858 → +0.663d_a rescored -0.824 → -0.071d_a mined -0.824d_b mined 0.858min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_-UiWTavWGiYtotal 53.0schain gain +2.4 dBseam step 1.4 dBcrossfades 150/100/150/150 ms
Script — 5 chunks, 4 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, normally alert, some disfluency, light breath
(elation, hope enthusiasm optimism · normal-paced, neutral tension, moderately variable, casual) (low mumble) Uh, it kept getting pushed further and further out, but we are very happy to announce that this is now totally live and up and anyone can go and download it. So the idea is this is not a research paper, this is just kind of a technical how-to of
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; somewhat unclear, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as elation, hope enthusiasm optimism; style: casual, conversational; average recording, quiet background; genuineness 3.3/6; vocal-burst blend 6.6/10; 12.6s, EN.
EN_-UiWTavWGiY_W000022 · in -22.1 dBFS · gain +2.1 dB · emolia-02480
(relief, hope enthusiasm optimism, contentment · normal-paced, slightly relaxed, fairly steady, casual) (ahem) Uhm, talk to me about this, the paper is up now, please go download it, (low mumble) uhm, get in contact with us, share your thoughts, questions, anything like that. So, this is just kind of an aside, but we're really happy to announce that we pushed to get it ready for BroCon and it is out and available now.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as relief, hope enthusiasm optimism, contentment; style: casual, monologue; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 4.0/10; 13.8s, EN.
EN_-UiWTavWGiY_W000024 · in -22.2 dBFS · gain +2.2 dB · emolia-02480
(confusion, embarrassment · normal-paced, slightly relaxed, fairly steady, casual) So, how do we do IR with Bro? (low mumble) Uhm, we are kind of an old school team. We don't have a sim really, except for Gmail. (low mumble)
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as confusion, embarrassment; style: casual, playful; good recording, quiet background; genuineness 3.3/6; vocal-burst blend 1.7/10; 7.5s, EN.
EN_-UiWTavWGiY_W000025 · in -19.9 dBFS · gain -0.1 dB · emolia-02480
(triumph, longing, elation · normal-paced, neutral tension, fairly steady, casual) We've, we've messed around with a lot of different things over the years. We've, we've had Splunk, we've had, uh, (low mumble) you know, the, the Elk stack, we've messed around with that, we've been looking at Spark, we've been looking at all this stuff, but ultimately our team often gets back to,
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as triumph, longing, elation; style: casual, conversational; average recording, quiet background; genuineness 4.3/6; vocal-burst blend 6.8/10; 13.0s, EN.
EN_-UiWTavWGiY_W000026 · in -20.9 dBFS · gain +0.9 dB · emolia-02480
(emotional numbness · brisk, slightly relaxed, fairly steady, monologue) And we pull all of our logs from most of our bros into like a central log repository and we, and then we've got some high powered crunching boxes to do that. So.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as emotional numbness; style: monologue, casual; good recording, no background noise; genuineness 1.6/6; vocal-burst blend 2.0/10; 6.7s, EN.
EN_-UiWTavWGiY_W000028 · in -18.5 dBFS · gain -1.5 dB · emolia-02480
Emotional Numbness(unconstrained axis: Hope Enthusiasm Optimism)identity −0.07 emotion 72 %   emotion__B1__T0.80__C0.25__INTERNAL · #14

This chain comes from the one-sided rule: only Emotional Numbness had to get where it was going, by at least 0.80. The other emotion was left completely free.

The chain starts with Emotional Numbness barely there — 0.13, lower than 87 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.84.

Nothing was asked of the other axis, and in fact Hope Enthusiasm Optimism drifts down from 0.89 to 0.38 (-0.51), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are +0.20, then +0.17, then +0.23, then +0.25 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.77 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.77 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 36 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.754 before conversion and 0.681 after — it fell by 0.073. Neighbour-to-neighbour the worst pair went 0.754 → 0.750. (The earlier render, with segment 1 left raw, scores 0.590 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.837 in the original and +0.601 after conversion — 72 % of the delta retained, which is most of it. On the other named axis, Hope Enthusiasm Optimism, -0.509 became -0.335.

Quality. Mean predicted overall quality across the segments went 2.87 → 2.98 (+0.11) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.754 → 0.681 -0.073identity cos neighbours 0.754 → 0.750d_b rescored +0.837 → +0.601d_a rescored -0.509 → -0.335d_a mined -0.509d_b mined 0.838min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_B00026_S03384total 34.5schain gain +1.8 dBseam step 1.3 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a child masculine voice · neutral-toned, neutral-bright, fairly smooth, good recording, no background noise, energised, slightly relaxed, clear
(measured, fairly steady, almost no disfluency, narration) In the summertime, ducks and geese migrate to the Arctic to build nests and raise their young. So if plants and animals can survive in the far north, what about people?
full caption & clip details
A child masculine voice; delivery is energised, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, almost no disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: narration, monologue; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.3/10; 12.1s, EN.
EN_B00026_S03384_W000023 · in -17.3 dBFS · gain -2.7 dB · emolia-00752
(fear, doubt, distress · normal-paced, fairly steady, no disfluency, playful) How would you stay warm during the cold, dark winters?
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as fear, doubt, distress; style: playful, formal; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 2.7/10; 3.6s, EN.
EN_B00026_S03384_W000024 · in -17.0 dBFS · gain -3.0 dB · emolia-00752
(fear, awe, doubt · measured, moderately variable, almost no disfluency, narration) How would you stay protected from the icy winds and snow storms? How would you find food?
full caption & clip details
An adult masculine voice; delivery is energised, measured, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, almost no disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as fear, awe, doubt; style: narration, storytelling; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 0.8/10; 6.2s, EN.
EN_B00026_S03384_W000025 · in -17.9 dBFS · gain -2.1 dB · emolia-00752
(normal-paced, fairly steady, almost no disfluency, narration) Before there were stores, fancy jackets, or electricity, these people survived in the frozen north.
full caption & clip details
A middle-aged masculine voice; delivery is energised, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: narration, formal; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.3/10; 7.8s, EN.
EN_B00026_S03384_W000026 · in -17.9 dBFS · gain -2.1 dB · emolia-00752
(emotional numbness · measured, fairly steady, almost no disfluency, narration) They built houses from driftwood, earth, whale bones, and snow.
full caption & clip details
An adult masculine voice; delivery is energised, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, almost no disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as emotional numbness; style: narration, monologue; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 2.4/10; 5.5s, EN.
EN_B00026_S03384_W000027 · in -18.0 dBFS · gain -2.0 dB · emolia-00752
Emotional Numbness(unconstrained axis: Astonishment Surprise)identity +0.31 emotion 98 %   emotion__B1__T0.80__C0.25__INTERNAL · #15

This chain comes from the one-sided rule: only Emotional Numbness had to get where it was going, by at least 0.80. The other emotion was left completely free.

The chain starts with Emotional Numbness essentially absent — 0.05, lower than 95 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.89.

Nothing was asked of the other axis, and in fact Astonishment Surprise drifts down from 0.84 to 0.14 (-0.71), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are +0.17, then +0.25, then +0.23, then +0.25 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the snippets clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 31 s · snippets

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was -0.075 before conversion and 0.239 after — it rose by 0.315. Neighbour-to-neighbour the worst pair went 0.037 → 0.465. (The earlier render, with segment 1 left raw, scores 0.311 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.921 in the original and +0.903 after conversion — 98 % of the delta retained, which is essentially all of it. On the other named axis, Astonishment Surprise, -0.704 became -0.677.

Quality. Mean predicted overall quality across the segments went 2.76 → 2.86 (+0.10) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 -0.075 → 0.239 +0.315identity cos neighbours 0.037 → 0.465d_b rescored +0.921 → +0.903d_a rescored -0.704 → -0.677d_a mined -0.706d_b mined 0.890min_cos_consec (site) —min_cos_anchor (site) —dataset snippetslang ?speaker batch2_part3_batch2_part3_total 30.0schain gain +2.1 dBseam step 1.4 dBcrossfades 100/100/100/100 ms
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, fairly smooth
(brisk, energised, neutral tension, storytelling) That driver, notably the only person on the bus, wearing a seatbelt.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, thin; clear, almost no disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; no dominant emotion; style: storytelling, dramatic; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 5.0/10; 4.5s.
batch2_part3_batch2_part3_chunk_100_1_87887 · in -26.8 dBFS · gain +6.8 dB · snippets-01045
(measured, normally alert, slightly relaxed, monologue) We've got to keep the students properly restrained in the event.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, thin; clear, almost no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, narration; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 0.7/10; 4.4s.
batch2_part3_batch2_part3_chunk_100_1_88004 · in -27.1 dBFS · gain +7.1 dB · snippets-01045
(brisk, normally alert, slightly relaxed, formal) While the National Highway Traffic Safety Administration says school buses are the most regulated vehicles on the road.
full caption & clip details
An adult feminine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; no dominant emotion; style: formal, dramatic; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 0.9/10; 6.0s.
batch2_part3_batch2_part3_chunk_100_1_88092 · in -26.0 dBFS · gain +6.0 dB · snippets-01045
(fatigue exhaustion · normal-paced, normally alert, slightly relaxed, monologue) Yeah with that I've daily network stats in the last 24 hours there are 648 transactions on the Swire network.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fatigue exhaustion; style: monologue; good recording, no background noise; genuineness 1.8/6; vocal-burst blend 1.6/10; 5.7s.
batch2_part3_batch2_part3_chunk_100_1_88203 · in -26.6 dBFS · gain +6.6 dB · snippets-01045
(emotional numbness, doubt, pain · measured, subdued, slightly relaxed, whispered) use a country's currency like that. (low mumble) Um, then you do have uh, (low mumble) you have the problem that they may for political reasons, (low mumble) um, just
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is neutral, neutral stance, neutral openness; reads as emotional numbness, doubt, pain; style: whispered, monologue; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 1.4/10; 9.8s.
batch2_part3_batch2_part3_chunk_100_1_88248 · in -27.4 dBFS · gain +7.4 dB · snippets-01045
Emotional Numbness(unconstrained axis: Concentration)identity −0.03 emotion 81 %   emotion__B1__T0.80__C0.25__INTERNAL · #16

This chain comes from the one-sided rule: only Emotional Numbness had to get where it was going, by at least 0.80. The other emotion was left completely free.

The chain starts with Emotional Numbness essentially absent — 0.01, lower than 99 % of clips in this corpus — and ends with it strongly present at 0.81, higher than 81 % of clips in this corpus. That is a total rise of 0.80.

Nothing was asked of the other axis, and in fact Concentration drifts down from 0.82 to 0.67 (-0.15), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are +0.22, then +0.22, then +0.16, then +0.21 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.65 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.65 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 55 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.639 before conversion and 0.608 after — it fell by 0.032. Neighbour-to-neighbour the worst pair went 0.722 → 0.571. (The earlier render, with segment 1 left raw, scores 0.517 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.805 in the original and +0.651 after conversion — 81 % of the delta retained, which is most of it. On the other named axis, Concentration, -0.151 became -0.138.

Quality. Mean predicted overall quality across the segments went 2.82 → 3.00 (+0.18) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.639 → 0.608 -0.032identity cos neighbours 0.722 → 0.571d_b rescored +0.805 → +0.651d_a rescored -0.151 → -0.138d_a mined -0.150d_b mined 0.805min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_oLoQX8RA6tItotal 54.2schain gain +2.0 dBseam step 0.3 dBcrossfades 150/100/150/100 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normally alert, slightly relaxed, moderate pitch range
(normal-paced, fairly steady, little disfluency, monologue) The next most important step is the Recode, or what in-game is called Accuracy.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 1.4/10; 4.4s, EN.
EN_oLoQX8RA6tI_W000022 · in -19.1 dBFS · gain -0.9 dB · emolia-00711
(normal-paced, fairly steady, almost no disfluency, monologue) As can be seen on screen, it has high recoil. It starts with vertical kick, then kicks to the right, returns to the center, only to kick upwards again.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 1.2/10; 7.8s, EN.
EN_oLoQX8RA6tI_W000023 · in -20.0 dBFS · gain +0.0 dB · emolia-00711
(concentration, triumph · normal-paced, fairly steady, some disfluency, monologue) Knowing this pattern makes predicting the recoil easier. Obviously using a grip improves the accuracy, but on this weapon it hardly reduces the recoil. You can see that its vertical kick is a little less, but it's not significant enough to make it a viable attachment. This means we're better off selecting another attachment.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, triumph; style: monologue, formal; average recording, quiet background; genuineness 0.5/6; vocal-burst blend 2.7/10; 18.3s, EN.
EN_oLoQX8RA6tI_W000024 · in -18.4 dBFS · gain -1.6 dB · emolia-00711
(concentration, interest, contentment · measured, steady, some disfluency, monologue) The hipfire spread isn't something to brag about either. We can't put it into a number, but on screen you can see the spread while standing, crouching and proning. While average for the assault rifles, it's smarter to aim down the sights. Keep in mind that this can change depending on our stance, as we can see the spread getting smaller when crouching or proning.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, minimal breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, interest, contentment; style: monologue, formal; average recording, quiet background; genuineness 0.0/6; vocal-burst blend 1.2/10; 19.2s, EN.
EN_oLoQX8RA6tI_W000025 · in -18.8 dBFS · gain -1.2 dB · emolia-00711
(normal-paced, steady, almost no disfluency, formal) Movement speed with the Rampart 17 is similar to the other assault rifles at 90%.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 1.0/10; 5.2s, EN.
EN_oLoQX8RA6tI_W000026 · in -16.7 dBFS · gain -3.3 dB · emolia-00711
Anger(unconstrained axis: Hope Enthusiasm Optimism)identity +0.01 emotion 76 %   emotion__B1__T0.80__C0.25__INTERNAL · #17

This chain comes from the one-sided rule: only Anger had to get where it was going, by at least 0.80. The other emotion was left completely free.

The chain starts with Anger essentially absent — 0.05, lower than 95 % of clips in this corpus — and ends with it strongly present at 0.86, higher than 86 % of clips in this corpus. That is a total rise of 0.82.

Nothing was asked of the other axis, and in fact Hope Enthusiasm Optimism barely moves at all, sitting near 0.88 throughout.

It takes 5 clips to get there. Clip to clip the moves are +0.23, then +0.20, then +0.19, then +0.19 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.88 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.88 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 44 s · de · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.828 before conversion and 0.835 after — it rose by 0.008. Neighbour-to-neighbour the worst pair went 0.727 → 0.831. (The earlier render, with segment 1 left raw, scores 0.686 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Anger moved +0.824 in the original and +0.629 after conversion — 76 % of the delta retained, which is most of it. On the other named axis, Hope Enthusiasm Optimism, -0.047 became -0.167.

Quality. Mean predicted overall quality across the segments went 2.93 → 3.20 (+0.27) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.828 → 0.835 +0.008identity cos neighbours 0.727 → 0.831d_b rescored +0.824 → +0.629d_a rescored -0.047 → -0.167d_a mined -0.048d_b mined 0.816min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_FSY0BnnloUAtotal 42.5schain gain +1.5 dBseam step 2.4 dBcrossfades 150/100/150/150 ms
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normally alert, slightly relaxed, some disfluency, average clarity
(brisk, moderately variable, wide pitch range, playful) Bei der Körperhaltung auf dem Motorrad gibt es so die ein oder anderen Missverständnisse, wie wir vielleicht in dem letzten Video schon erfahren haben. Da haben wir nämlich über die sogenannte Faustregel gesprochen.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; no dominant emotion; style: playful, conversational; good recording, quiet background; genuineness 2.2/6; vocal-burst blend 1.4/10; 11.6s, DE.
DE_FSY0BnnloUA_W000000 · in -12.6 dBFS · gain -7.4 dB · emolia-00151
(confusion, thankfulness gratitude · normal-paced, fairly steady, moderate pitch range, casual) Und heute widmen wir uns mal bisschen den Oberkörper und was es da so für, (low mumble) naja, ich sag mal jetzt nicht unbedingt Missverständnisse, aber ein paar so.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as confusion, thankfulness gratitude; style: casual, conversational; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 0.5/10; 9.8s, DE.
DE_FSY0BnnloUA_W000001 · in -12.3 dBFS · gain -7.7 dB · emolia-00151
(intoxication altered states of consciousness · normal-paced, fairly steady, moderate pitch range, casual) Falschmeinungen im Internet herumkursieren. Schaut es euch an. Ich bin gespannt, ob es der ein oder andere schon wusste. Bis gleich.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as intoxication altered states of consciousness; style: casual, conversational; good recording, quiet background; genuineness 3.3/6; vocal-burst blend 0.1/10; 7.0s, DE.
DE_FSY0BnnloUA_W000002 · in -13.6 dBFS · gain -6.4 dB · emolia-00151
(fear, confusion · normal-paced, fairly steady, moderate pitch range, casual) Man muss das ganze Thema Körperhaltung und allgemein die Haltung am Motorrad in mehrere Sequenzen aufschlüsseln, weil ich glaube nicht, dass es möglich ist, das Ganze in einem
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fear, confusion; style: casual, conversational; good recording, quiet background; genuineness 4.1/6; vocal-burst blend 0.7/10; 8.7s, DE.
DE_FSY0BnnloUA_W000003 · in -12.6 dBFS · gain -7.4 dB · emolia-00151
(brisk, fairly steady, moderate pitch range, dramatic) einigermaßen beizubringen. Und deswegen schlüssel ich das Thema so auf und deswegen unterhalten wir uns heute mal über den Oberkörper.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: dramatic, conversational; good recording, no background noise; genuineness 2.5/6; vocal-burst blend 0.9/10; 6.2s, DE.
DE_FSY0BnnloUA_W000004 · in -13.0 dBFS · gain -7.0 dB · emolia-00151
Relief(unconstrained axis: Confusion)identity +0.31 emotion 115 %   emotion__B1__T0.80__C0.25__INTERNAL · #18

This chain comes from the one-sided rule: only Relief had to get where it was going, by at least 0.80. The other emotion was left completely free.

The chain starts with Relief barely there — 0.19, lower than 81 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.80.

Nothing was asked of the other axis, and in fact Confusion drifts down from 0.98 to 0.77 (-0.21), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are +0.24, then +0.10, then +0.23, then +0.23 — an uneven climb, but always in the same direction.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.03 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.03 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 40 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.120 before conversion and 0.433 after — it rose by 0.313. Neighbour-to-neighbour the worst pair went 0.228 → 0.551. (The earlier render, with segment 1 left raw, scores 0.422 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Relief moved +0.790 in the original and +0.912 after conversion — 115 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Confusion, -0.215 became -0.201.

Quality. Mean predicted overall quality across the segments went 2.90 → 3.10 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.120 → 0.433 +0.313identity cos neighbours 0.228 → 0.551d_b rescored +0.790 → +0.912d_a rescored -0.215 → -0.201d_a mined -0.215d_b mined 0.803min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_ex5q3QUL95ototal 38.5schain gain +0.6 dBseam step 3.5 dBcrossfades 100/150/100/100 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a middle-aged masculine voice
(confusion, awe, contemplation · measured, normally alert, slightly relaxed, whispered) If someone says the Buddha has spoken spiritual truths, he slanders the Buddha due to his inability to understand what the Buddha teaches. Subhuti, as to speaking truth, no truth can be spoken, therefore it is called speaking truth.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; clear, almost no disfluency, fairly narrow pitch, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as confusion, awe, contemplation; style: whispered, narration; average recording, quiet background; genuineness 0.3/6; vocal-burst blend 0.9/10; 16.1s, EN.
EN_ex5q3QUL95o_W000278 · in -19.9 dBFS · gain -0.1 dB · emolia-02457
(measured, very low-energy, slightly relaxed, whispered) At that time, Subhuti, the wise elder, addressed the Buddha.
full caption & clip details
A child feminine voice; delivery is very low-energy, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: whispered, monologue; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 3.0/10; 4.5s, EN.
EN_ex5q3QUL95o_W000279 · in -18.7 dBFS · gain -1.3 dB · emolia-02457
(doubt, fear · measured, normally alert, slightly relaxed, whispered) Will there be living beings in the future who believe in this Sutra when they hear it?
full caption & clip details
An elderly masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is slightly cool, dark, fairly smooth, thin; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as doubt, fear; style: whispered, narration; average recording, no background noise; genuineness 1.4/6; vocal-burst blend 2.3/10; 4.4s, EN.
EN_ex5q3QUL95o_W000280 · in -15.7 dBFS · gain -4.3 dB · emolia-02457
(contemplation, awe, infatuation · slow, very low-energy, relaxed, whispered) The living beings to whom you refer Are neither living beings nor not living beings. Why, Subuti.
full caption & clip details
An elderly masculine voice; delivery is very low-energy, slow, relaxed, moderately variable; timbre is slightly warm, slightly dark, slightly rough, balanced body; average clarity, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly negative, slightly dominant, fairly guarded; reads as contemplation, awe, infatuation; style: whispered, storytelling; good recording, quiet background; genuineness 1.0/6; vocal-burst blend 1.5/10; 7.9s, EN.
EN_ex5q3QUL95o_W000282 · in -18.5 dBFS · gain -1.5 dB · emolia-02457
(relief, thankfulness gratitude, awe · measured, normally alert, slightly relaxed, whispered) Blessed Lord, when you attained complete enlightenment, did you feel in your mind that nothing had been acquired?
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as relief, thankfulness gratitude, awe; style: whispered, narration; good recording, no background noise; genuineness 1.3/6; vocal-burst blend 0.5/10; 6.2s, EN.
EN_ex5q3QUL95o_W000285 · in -19.3 dBFS · gain -0.7 dB · emolia-02457
Triumph(unconstrained axis: Fatigue Exhaustion)identity −0.10 emotion REVERSED   emotion__B1__T0.80__C0.25__INTERNAL · #19

This chain comes from the one-sided rule: only Triumph had to get where it was going, by at least 0.80. The other emotion was left completely free.

The chain starts with Triumph essentially absent — 0.04, lower than 96 % of clips in this corpus — and ends with it strongly present at 0.86, higher than 86 % of clips in this corpus. That is a total rise of 0.82.

Nothing was asked of the other axis, and in fact Fatigue Exhaustion drifts down from 0.98 to 0.19 (-0.79), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are +0.21, then +0.20, then +0.19, then +0.21 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.89 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.89 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 32 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.845 before conversion and 0.750 after — it fell by 0.096. Neighbour-to-neighbour the worst pair went 0.824 → 0.722. (The earlier render, with segment 1 left raw, scores 0.680 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

The emotional move did not survive. Re-scored end to end, Triumph moved +0.817 in the original and -0.400 after conversion — it changed direction. On this chain the corrected audio is not an improvement. On the other named axis, Fatigue Exhaustion, -0.789 became -0.849.

Quality. Mean predicted overall quality across the segments went 2.79 → 2.88 (+0.09) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.845 → 0.750 -0.096identity cos neighbours 0.824 → 0.722d_b rescored +0.817 → -0.400d_a rescored -0.789 → -0.849d_a mined -0.789d_b mined 0.817min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_B00027_S03560total 30.1schain gain +0.3 dBseam step 1.4 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: a child somewhat masculine voice
(fatigue exhaustion · normal-paced, energised, neutral tension, casual) And if we have more meeples than five on our card, we have to give it back in the middle and get points for the meeple.
full caption & clip details
A child somewhat masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as fatigue exhaustion; style: casual, playful; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 1.8/10; 6.9s, EN.
EN_B00027_S03560_W000005 · in -19.2 dBFS · gain -0.8 dB · emolia-00764
(intoxication altered states of consciousness · measured, normally alert, fully relaxed, casual) We take, if we take a card, we (low mumble) uhm, push the line in front and get a new card.
full caption & clip details
A child masculine voice; delivery is normally alert, measured, fully relaxed, moderately variable; timbre is slightly cool, slightly dark, slightly rough, thin; slurred, frequent disfluency, wide pitch range, normal breath; affect is neutral, neutral stance, neutral openness; reads as intoxication altered states of consciousness; style: casual, playful; below-average recording, quiet background; genuineness 3.0/6; vocal-burst blend 0.9/10; 7.1s, EN.
EN_B00027_S03560_W000006 · in -20.2 dBFS · gain +0.2 dB · emolia-00764
(normal-paced, normally alert, fully relaxed, casual) So the (ahem) (low mumble) expensive card is only every time the last one.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, fully relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 4.7/6; vocal-burst blend 2.6/10; 5.2s, EN.
EN_B00027_S03560_W000007 · in -23.2 dBFS · gain +3.2 dB · emolia-00764
(normal-paced, energised, slightly relaxed, playful) We take a card and we have, (ahem) uhm, seven different characters and seven different buildings.
full caption & clip details
A child masculine voice; delivery is energised, normal-paced, slightly relaxed, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; no dominant emotion; style: playful, casual; average recording, some background noise; genuineness 2.2/6; vocal-burst blend 1.0/10; 6.1s, EN.
EN_B00027_S03560_W000008 · in -19.9 dBFS · gain -0.1 dB · emolia-00764
(normal-paced, normally alert, slightly relaxed, casual) And for each building we have one character. So we take a character and...
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 1.1/10; 5.6s, EN.
EN_B00027_S03560_W000009 · in -20.6 dBFS · gain +0.6 dB · emolia-00764
Thankfulness Gratitude(unconstrained axis: Concentration)identity +0.25 emotion 113 %   emotion__B1__T0.80__C0.25__INTERNAL · #20

This chain comes from the one-sided rule: only Thankfulness Gratitude had to get where it was going, by at least 0.80. The other emotion was left completely free.

The chain starts with Thankfulness Gratitude barely there — 0.17, lower than 83 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.83.

Nothing was asked of the other axis, and in fact Concentration drifts down from 0.93 to 0.16 (-0.77), which the rule did not require.

It takes 5 clips to get there. Clip to clip the moves are +0.16, then +0.24, then +0.23, then +0.19 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.24 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.24 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 27 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.090 before conversion and 0.336 after — it rose by 0.246. Neighbour-to-neighbour the worst pair went 0.075 → 0.343. (The earlier render, with segment 1 left raw, scores 0.283 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Thankfulness Gratitude moved +0.850 in the original and +0.961 after conversion — 113 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Concentration, -0.771 became -0.730.

Quality. Mean predicted overall quality across the segments went 2.54 → 2.82 (+0.28) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.090 → 0.336 +0.246identity cos neighbours 0.075 → 0.343d_b rescored +0.850 → +0.961d_a rescored -0.771 → -0.730d_a mined -0.771d_b mined 0.830min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_tAW9zt_jqmutotal 25.4schain gain +3.3 dBseam step 1.6 dBcrossfades 100/150/100/150 ms
Script — 5 chunks, 2 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, fairly smooth, balanced body, normally alert, slightly relaxed, fairly steady, moderate pitch range, light breath
(concentration, interest · normal-paced, some disfluency, average clarity, casual) (low mumble) Uhm, we're planning, this is, (ahem) uh, you know, kind of a quasi longitudinal effort, so we'll try to keep a certain number of the fields we used in the original study stable.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as concentration, interest; style: casual; average recording, quiet background; genuineness 2.4/6; vocal-burst blend 1.4/10; 9.3s, EN.
EN_tAW9zt_jqmu_W000154 · in -17.7 dBFS · gain -2.3 dB · emolia-02308
(normal-paced, frequent disfluency, average clarity, casual) But, (ahem) uh, we especially are interested to hear if there's certain kinds of
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; no dominant emotion; style: casual, monologue; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 1.2/10; 4.3s, EN.
EN_tAW9zt_jqmu_W000155 · in -22.8 dBFS · gain +2.8 dB · emolia-02308
(normal-paced, no disfluency, clear, casual) Data points or metrics you'd like to see tracked over time in this area.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 2.4/10; 3.5s, EN.
EN_tAW9zt_jqmu_W000156 · in -20.3 dBFS · gain +0.3 dB · emolia-02308
(brisk, almost no disfluency, clear, playful) But again, if you want to ask us any other questions about what we presented that is also welcome to you, do not feel beholden to these questions at all.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; no dominant emotion; style: playful, dramatic; good recording, quiet background; genuineness 2.2/6; vocal-burst blend 1.8/10; 5.5s, EN.
EN_tAW9zt_jqmu_W000157 · in -19.0 dBFS · gain -1.0 dB · emolia-02308
(thankfulness gratitude, relief, embarrassment · normal-paced, some disfluency, average clarity, casual) Good afternoon, thank you for your presentation. I was not able to
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as thankfulness gratitude, relief, embarrassment; style: casual, conversational; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 1.0/10; 3.5s, EN.
EN_tAW9zt_jqmu_W000158 · in -19.3 dBFS · gain -0.7 dB · emolia-02308