Manifest tier. proxy_spearman, rule PXR, T=0.25, step cap 0.25. Population 338,690 chains (3,756 h) over 6 corpora. The SHAREABLE variant of this tier (podcast and evasnippets excluded) holds 271,942.
How to read a Script. Each chunk is one line:
a short tag of what the models heard in that clip, then the words spoken.
(underlined, plain ·
delivery, style) — the tag before the words. Emotions first, then how
it is delivered. Underlined descriptors are the ones that change across this chain
— anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words
are a real non-speech sound, printed where it happens.
The full generated caption for any clip is under “full caption & clip
details”. Its perceived-gender and background-noise clauses were re-rendered from
the numeric buckets, because the versions stored in the corpus index had those two
ladders running backwards.
The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.
The tags and captions describe the ORIGINAL clips — they are the
corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the
converted audio are the ones printed in each card's own paragraph and chips.
This chain comes from the proxy rule: the same two-sided test as above, but because Triumph is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Triumph below average — 0.32, lower than 68 % of clips in this corpus — and ends with it strongly present at 0.86, higher than 86 % of clips in this corpus. That is a total rise of 0.54.
At the same time Pain goes the other way, from 0.72 (higher than 72 % of clips in this corpus) to 0.35 (lower than 65 % of clips in this corpus), a change of -0.37. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.13, then +0.24, then +0.17 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.87 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.92 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.87. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 30 s · ko · emolia
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.878 before conversion and 0.827 after — it fell by 0.051. Neighbour-to-neighbour the worst pair went 0.852 → 0.859. (The earlier render, with segment 1 left raw, scores 0.799 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
The emotional move did not survive. Re-scored end to end, Triumph moved +0.541 in the original and -0.134 after conversion — it changed direction. On this chain the corrected audio is not an improvement. On the other named axis, Pain, -0.369 became -0.173.
Quality. Mean predicted overall quality across the segments went 3.09 → 3.24 (+0.15) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.878 → 0.827-0.051identity cos neighbours 0.852 → 0.859d_b rescored +0.541 → -0.134d_a rescored -0.369 → -0.173d_a mined -0.369d_b mined 0.541min_cos_consec (site) 0.9248min_cos_anchor (site) 0.8738dataset emolialang kospeaker KO_yETieOt82j4total 29.3schain gain +2.1 dBseam step 0.8 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, no background noise, normally alert, slightly relaxed, fairly steady
(measured, almost no disfluency, clear, authoritative)자, 이제 마지막 세번째로 가보겠습니다. 미국 배당자산 이태보풍산 투자입니다.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: authoritative, formal; good recording, no background noise; genuineness 1.3/6; vocal-burst blend 0.0/10; 6.4s, KO.
KO_yETieOt82j4_W000062 · in -19.7 dBFS · gain -0.3 dB · emolia-03140
(measured, some disfluency, somewhat unclear, monologue)자, 월 배당 목돈 운용인데 이것은 저희 바인 투자 자문에서 투자 자문 해서 삼성 정권을 통해서 가입 운용하는 건데, 지난해
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, didactic; average recording, no background noise; genuineness 2.7/6; vocal-burst blend 0.1/10; 10.0s, KO.
KO_yETieOt82j4_W000063 · in -19.9 dBFS · gain -0.1 dB · emolia-03140
(normal-paced, little disfluency, average clarity, monologue)연 4.93% 나왔습니다. 2021년 기준 배당이 4.93% 나왔는데, 앞서 말씀드린 것처럼 이것을 왜,
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, authoritative; good recording, no background noise; genuineness 1.6/6; vocal-burst blend 0.2/10; 8.7s, KO.
KO_yETieOt82j4_W000064 · in -21.8 dBFS · gain +1.8 dB · emolia-03140
(normal-paced, little disfluency, average clarity, formal)배당 주식이라고 하지 않고 배당 자산이라고 표현하냐, 앞서도 설명드렸지만,
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 1.2/6; vocal-burst blend 0.5/10; 4.9s, KO.
KO_yETieOt82j4_W000065 · in -20.1 dBFS · gain +0.1 dB · emolia-03140
This chain comes from the proxy rule: the same two-sided test as above, but because Bitterness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Bitterness clearly present — 0.65, higher than 65 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.34.
At the same time Doubt goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.73 (higher than 73 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.23, then +0.09, then +0.02 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.47 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.54 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.47, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 111 s · pt · podcast
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.490 before conversion and 0.831 after — it rose by 0.341. Neighbour-to-neighbour the worst pair went 0.521 → 0.831. (The earlier render, with segment 1 left raw, scores 0.603 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Bitterness moved +0.330 in the original and +0.366 after conversion — 111 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Doubt, -0.258 became -0.621.
Quality. Mean predicted overall quality across the segments went 2.69 → 3.23 (+0.54) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.490 → 0.831+0.341identity cos neighbours 0.521 → 0.831d_b rescored +0.330 → +0.366d_a rescored -0.258 → -0.621d_a mined -0.262d_b mined 0.345min_cos_consec (site) 0.5390min_cos_anchor (site) 0.4742dataset podcastlang ptspeaker 885930total 110.3schain gain +1.8 dBseam step 1.2 dBcrossfades 150/150/150 ms
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: a middle-aged masculine voice · slightly rough, quiet background, light breath
(doubt, contemplation, relief · measured, subdued, relaxed, monologue)Yeah, you follow the autoconhecimento is important, because observing, can you, okay, so there are people who don't (low mumble) have this comprehension, to autoconhecer. E aí gera um desafio to
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as doubt, contemplation, relief; style: monologue, whispered; average recording, quiet background; genuineness 2.7/6; vocal-burst blend 3.6/10; 27.0s, PT.
885930_00122768 · in -18.6 dBFS · gain -1.4 dB · podcast-01243
(pride, hope enthusiasm optimism, awe·normal-paced, normally alert, slightly relaxed, monologue)sabe que você está se desafiando no dia a dia, ou o quanto você está vivendo no automatic. Vive no automatically, colocando aqui um conceito budista, you could colour the 10 status of it, is vive those prime status of it, which are inferno, fome, animalidade, ira, tranquility and alegria. When you decide to vive in quality one,
full caption & clip details
A child masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as pride, hope enthusiasm optimism, awe; style: monologue, cartoonish; average recording, quiet background; genuineness 1.8/6; vocal-burst blend 4.7/10; 28.6s, PT.
885930_00126400 · in -14.5 dBFS · gain -5.5 dB · podcast-05107
(bitterness, disgust, interest·brisk, energised, neutral tension, dramatic)Você está vivendo no automático, geralmente. Por quê? Porque o ambiente molda as suas circunstâncias, ele molda a sua reatividade, ele dá um looping, ele mantém a sua reatividade, porque você reage, né? Manifestando esses mesmos estados, que vão fazer com que o ambiente ainda seja mais forte e gere uma reatividade mais intensa em você e você não saia deles. É igual quando você está naquela praia que tem aquelas ondas fortes, que você está na
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as bitterness, disgust, interest; style: dramatic, monologue; below-average recording, quiet background; genuineness 3.3/6; vocal-burst blend 10.0/10; 29.9s, PT.
885930_00129256 · in -14.2 dBFS · gain -5.8 dB · podcast-05111
(bitterness, sourness, interest ·fast, energised, neutral tension, casual)nesse ciclo desses seis estados de vida, né? Que inclusive em outras linhas budistas é muito forte esse conceito chamado (low mumble) a roda do samsara, que são esses baixos estados de vida que mantém o quê? Os desejos mundanos, que faz com que você realmente se deixe levar. Mas quando você começa a restar o Nami Ohorengyekyo, você já sai deles, porque você está procurando alguma coisa.
full caption & clip details
A young adult masculine voice; delivery is energised, fast, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, slightly thin; somewhat unclear, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as bitterness, sourness, interest; style: casual, dramatic; below-average recording, quiet background; genuineness 4.8/6; vocal-burst blend 10.0/10; 25.3s, PT.
885930_00135352 · in -15.6 dBFS · gain -4.4 dB · podcast-05104
This chain comes from the proxy rule: the same two-sided test as above, but because Embarrassment is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Embarrassment clearly present — 0.68, higher than 68 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.26.
At the same time Teasing goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.36 (lower than 64 % of clips in this corpus), a change of -0.63. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.19, then +0.08 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? Not directly measured. What does exist is a timbre similarity of 0.90 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.
Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.90 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 16 s · en · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.735 before conversion and 0.574 after — it fell by 0.162. Neighbour-to-neighbour the worst pair went 0.796 → 0.686. (The earlier render, with segment 1 left raw, scores 0.503 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Embarrassment moved +0.262 in the original and +0.190 after conversion — 72 % of the delta retained, which is most of it. On the other named axis, Teasing, -0.626 became -0.197.
Quality. Mean predicted overall quality across the segments went 2.55 → 2.77 (+0.22) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.735 → 0.574-0.162identity cos neighbours 0.796 → 0.686d_b rescored +0.262 → +0.190d_a rescored -0.626 → -0.197d_a mined -0.626d_b mined 0.261min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_un79XT7KcIktotal 14.9schain gain +1.7 dBseam step 1.0 dBcrossfades 150/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a child feminine voice · slightly bright, below-average recording, moderately variable, some disfluency, wide pitch range
(teasing, astonishment surprise, affection · brisk, energised, neutral tension, casual)I'm sure some of you probably have thought that too. Lupa's in the ground. Check it out.
full caption & clip details
A child feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, slightly rough, thin; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as teasing, astonishment surprise, affection; style: casual, conversational; below-average recording, some background noise; genuineness 4.3/6; vocal-burst blend 3.6/10; 4.6s, EN.
EN_un79XT7KcIk_W000043 · in -16.4 dBFS · gain -3.6 dB · emolia-02308
(amusement· brisk, energised, neutral tension, casual)We've got four plants of loofah there. Well, I've got more because on I think two of them.
full caption & clip details
A child feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, slightly bright, slightly rough, thin; somewhat unclear, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as amusement; style: casual, conversational; below-average recording, some background noise; genuineness 4.2/6; vocal-burst blend 4.6/10; 5.7s, EN.
EN_un79XT7KcIk_W000044 · in -15.9 dBFS · gain -4.1 dB · emolia-02308
(embarrassment, astonishment surprise, amusement ·normal-paced, very low-energy, relaxed, conversational)There is, ah, sit down. I think two of (childlike giggle) them, there's two plants in one hole.
full caption & clip details
A child feminine voice; delivery is very low-energy, normal-paced, relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; somewhat unclear, some disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as embarrassment, astonishment surprise, amusement; style: conversational, casual; below-average recording, quiet background; genuineness 4.2/6; vocal-burst blend 3.9/10; 5.0s, EN.
EN_un79XT7KcIk_W000045 · in -15.0 dBFS · gain -5.0 dB · emolia-02308
Intoxication Altered States of Consciousness ↓ / Emotional Numbness ↑identity −0.14emotion 85 % proxy_spearman__PXR__T0.25__C0.25__INTERNAL · #4
This chain comes from the proxy rule: the same two-sided test as above, but because Emotional Numbness is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Emotional Numbness around average — 0.54, higher than 54 % of clips in this corpus — and ends with it at the very top of the corpus at 0.91, higher than 91 % of clips in this corpus. That is a total rise of 0.38.
At the same time Intoxication Altered States of Consciousness goes the other way, from 0.70 (higher than 70 % of clips in this corpus) to 0.36 (lower than 64 % of clips in this corpus), a change of -0.34. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.11, then +0.24, then +0.03 — a plateau around step 3, where it barely moves.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.80 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.80 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.80, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 27 s · en · emolia
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.795 before conversion and 0.653 after — it fell by 0.142. Neighbour-to-neighbour the worst pair went 0.801 → 0.653. (The earlier render, with segment 1 left raw, scores 0.611 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.376 in the original and +0.321 after conversion — 85 % of the delta retained, which is most of it. On the other named axis, Intoxication Altered States of Consciousness, -0.341 became -0.072.
Quality. Mean predicted overall quality across the segments went 2.58 → 2.81 (+0.23) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.795 → 0.653-0.142identity cos neighbours 0.801 → 0.653d_b rescored +0.376 → +0.321d_a rescored -0.341 → -0.072d_a mined -0.341d_b mined 0.376min_cos_consec (site) 0.7963min_cos_anchor (site) 0.7957dataset emolialang enspeaker EN_B00067_S07520total 26.0schain gain +2.0 dBseam step 0.3 dBcrossfades 150/150/150 ms
Script — 4 chunks, 4 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · slightly cool, quiet background, frequent disfluency
(measured, normally alert, slightly relaxed, ASMR)Now I have this (low mumble) table here with the Latin terms, the English terms.
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, fairly smooth, slightly thin; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; no dominant emotion; style: ASMR, casual; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 0.3/10; 5.4s, EN.
EN_B00067_S07520_W000003 · in -19.9 dBFS · gain -0.1 dB · emolia-01521
(slow, very low-energy, slightly relaxed, whispered)And (low mumble) Romanian terms from five sources with references to the pages.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, slow, slightly relaxed, steady; timbre is slightly cool, slightly dark, fairly smooth, slightly thin; average clarity, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly positive, slightly submissive, neutral openness; no dominant emotion; style: whispered, monologue; average recording, quiet background; genuineness 2.4/6; vocal-burst blend 0.3/10; 6.0s, EN.
EN_B00067_S07520_W000004 · in -20.3 dBFS · gain +0.3 dB · emolia-01521
(concentration· slow, very low-energy, relaxed, monologue)Uh, (low mumble) these were the matches done by the algorithm based, based on, (low mumble) uhm, the Google translation.
full caption & clip details
A middle-aged feminine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is slightly cool, slightly dark, slightly rough, thin; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly positive, slightly submissive, neutral openness; reads as concentration; style: monologue, whispered; below-average recording, quiet background; genuineness 4.3/6; vocal-burst blend 0.4/10; 7.5s, EN.
EN_B00067_S07520_W000005 · in -20.0 dBFS · gain -0.0 dB · emolia-01521
(emotional numbness, pain· slow, very low-energy, relaxed, monologue)And then (low mumble) downloaded as in csv and uploaded in R. So the connections between
full caption & clip details
A young adult feminine voice; delivery is very low-energy, slow, relaxed, steady; timbre is slightly cool, neutral-bright, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly positive, submissive, neutral openness; reads as emotional numbness, pain; style: monologue, casual; below-average recording, quiet background; genuineness 3.7/6; vocal-burst blend 0.0/10; 7.7s, EN.
EN_B00067_S07520_W000006 · in -21.0 dBFS · gain +1.0 dB · emolia-01521
This chain comes from the proxy rule: the same two-sided test as above, but because Affection is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Affection clearly present — 0.60, higher than 60 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.38.
At the same time Pride goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.73 (higher than 73 % of clips in this corpus), a change of -0.27. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.20, then +0.17 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.87 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.89 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.87. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 38 s · en · podcast
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.782 before conversion and 0.747 after — it fell by 0.035. Neighbour-to-neighbour the worst pair went 0.699 → 0.672. (The earlier render, with segment 1 left raw, scores 0.632 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Affection moved +0.377 in the original and +0.336 after conversion — 89 % of the delta retained, which is most of it. On the other named axis, Pride, -0.267 became -0.261.
Quality. Mean predicted overall quality across the segments went 2.91 → 3.12 (+0.21) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.782 → 0.747-0.035identity cos neighbours 0.699 → 0.672d_b rescored +0.377 → +0.336d_a rescored -0.267 → -0.261d_a mined -0.268d_b mined 0.376min_cos_consec (site) 0.8873min_cos_anchor (site) 0.8726dataset podcastlang enspeaker 111801total 37.4schain gain +2.2 dBseam step 1.4 dBcrossfades 150/100 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, balanced body, quiet background, brisk, energised, moderately variable, some disfluency, average clarity
(pride, sexual lust, shame · slightly tense, light breath, casual, conversational)but the satisfaction and gratification you get to knowing that I'm betting on myself. When you bet on yourself, you set yourself to a standard.
full caption & clip details
An adult masculine voice; delivery is energised, brisk, slightly tense, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, fairly guarded; reads as pride, sexual lust, shame; style: casual, conversational; good recording, quiet background; genuineness 2.9/6; vocal-burst blend 4.1/10; 6.8s, EN.
111801_00038408 · in -31.7 dBFS · gain +11.7 dB · podcast-04632
(disappointment, impatience and irritability, fatigue exhaustion·neutral tension, normal breath, casual, conversational)You set yourself to accountability of bruh, because at the end of the day, if you don't give, if you don't give yourself 110% every day, why are you even doing this? Go back to the nine to five. Go back to it. I wake up every day ready to conquer the day. I can't, I can't count it in in two week paychecks. I gotta count it every day. Somebody who said, well, when do you take a day off? You don't take no day off when you out here grinding, trying
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is mildly negative, slightly dominant, fairly guarded; reads as disappointment, impatience and irritability, fatigue exhaustion; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 4.3/6; vocal-burst blend 7.3/10; 24.9s, EN.
111801_00039104 · in -30.8 dBFS · gain +10.8 dB · podcast-06249
(affection, hope enthusiasm optimism, elation·slightly tense, light breath, casual, conversational)no days off. At every point in time, you should be watching some video, learning something. You should be talking to somebody. You should be calling
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, slightly tense, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, fairly guarded; reads as affection, hope enthusiasm optimism, elation; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 4.1/6; vocal-burst blend 7.5/10; 6.0s, EN.
111801_00041712 · in -28.7 dBFS · gain +8.7 dB · podcast-04614
This chain comes from the proxy rule: the same two-sided test as above, but because Infatuation is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Infatuation barely there — 0.25, lower than 75 % of clips in this corpus — and ends with it clearly present at 0.73, higher than 74 % of clips in this corpus. That is a total rise of 0.49.
At the same time Emotional Numbness goes the other way, from 0.87 (higher than 87 % of clips in this corpus) to 0.55 (higher than 55 % of clips in this corpus), a change of -0.33. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.14, then +0.21, then +0.14 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.94 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.95 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.94), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 38 s · en · emolia
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.906 before conversion and 0.834 after — it fell by 0.071. Neighbour-to-neighbour the worst pair went 0.926 → 0.880. (The earlier render, with segment 1 left raw, scores 0.802 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Infatuation moved +0.487 in the original and +0.420 after conversion — 86 % of the delta retained, which is most of it. On the other named axis, Emotional Numbness, -0.322 became -0.877.
Quality. Mean predicted overall quality across the segments went 3.08 → 3.16 (+0.08) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.906 → 0.834-0.071identity cos neighbours 0.926 → 0.880d_b rescored +0.487 → +0.420d_a rescored -0.322 → -0.877d_a mined -0.325d_b mined 0.487min_cos_consec (site) 0.9470min_cos_anchor (site) 0.9412dataset emolialang enspeaker EN_wClDFsmKjV8total 37.2schain gain +1.3 dBseam step 0.5 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, good recording, no background noise, normally alert, slightly relaxed, steady
(measured, fairly narrow pitch, formal, monologue)These have been collected by a robotic, "'vacuum cleaner", and examined, leading to improved estimates of their flux and mass distribution
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 11.1s, EN.
EN_wClDFsmKjV8_W000241 · in -16.7 dBFS · gain -3.3 dB · emolia-01046
(relief· measured, moderate pitch range, formal, monologue)The well is not an ice core, but the age of the ice that was melted is known, so the
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as relief; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 9.1s, EN.
EN_wClDFsmKjV8_W000242 · in -15.9 dBFS · gain -4.1 dB · emolia-01046
(normal-paced, moderate pitch range, formal, monologue)The well becomes about 10 metres deeper each year, so micrometeorites collected in a given year are about 100 years older than those from the previous year
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 11.1s, EN.
EN_wClDFsmKjV8_W000243 · in -15.3 dBFS · gain -4.7 dB · emolia-01046
(measured, moderate pitch range, formal, monologue)Pollen, an important component of sediment cores, can also be found in ice cores
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 6.5s, EN.
EN_wClDFsmKjV8_W000244 · in -15.7 dBFS · gain -4.3 dB · emolia-01046
This chain comes from the proxy rule: the same two-sided test as above, but because Concentration is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Concentration around average — 0.47, lower than 53 % of clips in this corpus — and ends with it clearly present at 0.74, higher than 74 % of clips in this corpus. That is a total rise of 0.27.
At the same time Emotional Numbness goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.63 (higher than 63 % of clips in this corpus), a change of -0.36. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.20, then +0.07 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.92 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.92 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.92), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 28 s · zh · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.898 before conversion and 0.917 after — it rose by 0.018. Neighbour-to-neighbour the worst pair went 0.891 → 0.887. (The earlier render, with segment 1 left raw, scores 0.882 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.271 in the original and +0.186 after conversion — 69 % of the delta retained. On the other named axis, Emotional Numbness, -0.357 became -0.165.
Quality. Mean predicted overall quality across the segments went 3.08 → 3.23 (+0.15) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.898 → 0.917+0.018identity cos neighbours 0.891 → 0.887d_b rescored +0.271 → +0.186d_a rescored -0.357 → -0.165d_a mined -0.358d_b mined 0.269min_cos_consec (site) 0.9186min_cos_anchor (site) 0.9215dataset emolialang zhspeaker ZH_B00064_S02495total 27.3schain gain +2.5 dBseam step 1.1 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, no background noise, normally alert, slightly relaxed, fairly steady
(emotional numbness, sexual lust, contempt · normal-paced, monologue, didactic)可就在我的刀即将砍到他脖子的时候,他却突然转过脑袋。对我诡异的一笑。
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, sexual lust, contempt; style: monologue, didactic; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 1.6/10; 7.1s, ZH.
ZH_B00064_S02495_W000043 · in -25.2 dBFS · gain +5.2 dB · emolia-03917
This chain comes from the proxy rule: the same two-sided test as above, but because Jealousy and Envy is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Jealousy and Envy below average — 0.39, lower than 61 % of clips in this corpus — and ends with it clearly present at 0.67, higher than 67 % of clips in this corpus. That is a total rise of 0.28.
At the same time Relief goes the other way, from 0.68 (higher than 68 % of clips in this corpus) to 0.24 (lower than 76 % of clips in this corpus), a change of -0.44. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.13, then +0.15 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.94 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.94 against each other.
Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.94), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 12 s · en · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.910 before conversion and 0.807 after — it fell by 0.103. Neighbour-to-neighbour the worst pair went 0.940 → 0.838. (The earlier render, with segment 1 left raw, scores 0.679 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
The emotional move did not survive. Re-scored end to end, Jealousy and Envy moved +0.276 in the original and -0.042 after conversion — it changed direction. On this chain the corrected audio is not an improvement. On the other named axis, Relief, -0.438 became -0.549.
Quality. Mean predicted overall quality across the segments went 2.66 → 2.79 (+0.14) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.910 → 0.807-0.103identity cos neighbours 0.940 → 0.838d_b rescored +0.276 → -0.042d_a rescored -0.438 → -0.549d_a mined -0.438d_b mined 0.276min_cos_consec (site) 0.9380min_cos_anchor (site) 0.9380dataset emolialang enspeaker EN_hSl2zuEMg0Mtotal 11.6schain gain +1.0 dBseam step 0.8 dBcrossfades 100/100 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(fairly steady, formal, authoritative)Run-up height, port of Ofunato area 24 meters 79 feet
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 1.0/10; 4.9s, EN.
EN_hSl2zuEMg0M_W000162 · in -13.7 dBFS · gain -6.3 dB · emolia-00629
(steady, formal, authoritative)Fishery port of Onagawa 15 metres, 49 feet
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 2.2/10; 3.9s, EN.
EN_hSl2zuEMg0M_W000163 · in -14.4 dBFS · gain -5.6 dB · emolia-00629
(fairly steady, formal, casual)Port of Ishinomaki 5 metres
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, casual; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 2.6/10; 3.2s, EN.
EN_hSl2zuEMg0M_W000164 · in -15.9 dBFS · gain -4.1 dB · emolia-00629
This chain comes from the proxy rule: the same two-sided test as above, but because Pain is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Pain below average — 0.35, lower than 65 % of clips in this corpus — and ends with it clearly present at 0.72, higher than 72 % of clips in this corpus. That is a total rise of 0.37.
At the same time Infatuation goes the other way, from 0.84 (higher than 84 % of clips in this corpus) to 0.45 (lower than 55 % of clips in this corpus), a change of -0.40. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.20, then +0.17 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.88 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.88 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.88. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 19 s · zh · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.840 before conversion and 0.835 after — it fell by 0.006. Neighbour-to-neighbour the worst pair went 0.856 → 0.835. (The earlier render, with segment 1 left raw, scores 0.794 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Pain moved +0.369 in the original and +0.369 after conversion — 100 % of the delta retained, which is essentially all of it. On the other named axis, Infatuation, -0.396 became -0.253.
Quality. Mean predicted overall quality across the segments went 3.03 → 3.09 (+0.06) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.840 → 0.835-0.006identity cos neighbours 0.856 → 0.835d_b rescored +0.369 → +0.369d_a rescored -0.396 → -0.253d_a mined -0.396d_b mined 0.369min_cos_consec (site) 0.8846min_cos_anchor (site) 0.8786dataset emolialang zhspeaker ZH_B00080_S07553total 18.5schain gain +2.6 dBseam step 1.4 dBcrossfades 100/100 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed
This chain comes from the proxy rule: the same two-sided test as above, but because Anger is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Anger clearly present — 0.73, higher than 73 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.27.
At the same time Astonishment Surprise goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.52 (higher than 52 % of clips in this corpus), a change of -0.48. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.19, then -0.15, then +0.23 — not a clean run: step 2 moves back the other way by 0.15 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.77 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.81 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.77, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 30 s · en · emolia
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.722 before conversion and 0.775 after — it rose by 0.052. Neighbour-to-neighbour the worst pair went 0.806 → 0.778. (The earlier render, with segment 1 left raw, scores 0.657 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Anger moved +0.270 in the original and +0.295 after conversion — 109 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Astonishment Surprise, -0.640 became -0.639.
Quality. Mean predicted overall quality across the segments went 2.81 → 3.08 (+0.27) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.722 → 0.775+0.052identity cos neighbours 0.806 → 0.778d_b rescored +0.270 → +0.295d_a rescored -0.640 → -0.639d_a mined -0.475d_b mined 0.270min_cos_consec (site) 0.8132min_cos_anchor (site) 0.7666dataset emolialang enspeaker EN_YJG7hNQaSQktotal 28.4schain gain +3.4 dBseam step 0.6 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · measured, normally alert, neutral tension, fairly steady
(astonishment surprise, fatigue exhaustion, awe · frequent disfluency, somewhat unclear, moderate pitch range, casual)My gosh, I'm telling you, we're in a season right now where God is saying, I want
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, neutral tension, fairly steady; timbre is neutral-toned, dark, slightly rough, thin; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is mildly negative, slightly dominant, neutral openness; reads as astonishment surprise, fatigue exhaustion, awe; style: casual, monologue; average recording, some background noise; genuineness 2.8/6; vocal-burst blend 2.4/10; 5.0s, EN.
EN_YJG7hNQaSQk_W000151 · in -12.9 dBFS · gain -7.1 dB · emolia-01222
(hope enthusiasm optimism, interest, fear·some disfluency, somewhat unclear, fairly narrow pitch, monologue)To release rewards. Listen, there's a promotion in the prophetic. I'm telling you right now. Who am I talking to? The Lord is promoting you. He's upgrading you. Listen.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, neutral tension, fairly steady; timbre is slightly cool, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, audible breath; affect is neutral, slightly dominant, slightly guarded; reads as hope enthusiasm optimism, interest, fear; style: monologue, authoritative; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 2.3/10; 8.5s, EN.
EN_YJG7hNQaSQk_W000152 · in -14.9 dBFS · gain -5.1 dB · emolia-01222
(fear, pain·frequent disfluency, average clarity, moderate pitch range, casual)Everybody watching right now, I know that you are in a sense a part of the remnant because
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, neutral tension, fairly steady; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as fear, pain; style: casual, monologue; average recording, quiet background; genuineness 2.8/6; vocal-burst blend 0.6/10; 6.3s, EN.
EN_YJG7hNQaSQk_W000153 · in -13.2 dBFS · gain -6.8 dB · emolia-01222
(anger, malevolence malice, hope enthusiasm optimism·some disfluency, somewhat unclear, moderate pitch range, casual)You're really hungry for the things of God. You're not afraid of COVID, okay? You love righteousness. You're pro-life. You're pro-Bible. You're pro-Israel, alright?
full caption & clip details
An elderly masculine voice; delivery is normally alert, measured, neutral tension, fairly steady; timbre is slightly cool, neutral-bright, slightly rough, thin; somewhat unclear, some disfluency, moderate pitch range, audible breath; affect is neutral, slightly dominant, fairly guarded; reads as anger, malevolence malice, hope enthusiasm optimism; style: casual, monologue; below-average recording, quiet background; genuineness 2.7/6; vocal-burst blend 4.2/10; 9.2s, EN.
EN_YJG7hNQaSQk_W000154 · in -13.3 dBFS · gain -6.7 dB · emolia-01222
This chain comes from the proxy rule: the same two-sided test as above, but because Malevolence Malice is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Malevolence Malice around average — 0.50, right about the corpus median — and ends with it at the very top of the corpus at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.42.
At the same time Confusion goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.64 (higher than 64 % of clips in this corpus), a change of -0.33. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.24, then +0.07, then +0.11 — most of the change happening immediately, then levelling off.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.73 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.73 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.73, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 37 s · zh · emolia
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.760 before conversion and 0.834 after — it rose by 0.074. Neighbour-to-neighbour the worst pair went 0.760 → 0.817. (The earlier render, with segment 1 left raw, scores 0.695 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Malevolence Malice moved +0.423 in the original and +0.029 after conversion — 7 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Confusion, -0.328 became -0.064.
Quality. Mean predicted overall quality across the segments went 2.77 → 3.20 (+0.43) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.760 → 0.834+0.074identity cos neighbours 0.760 → 0.817d_b rescored +0.423 → +0.029d_a rescored -0.328 → -0.064d_a mined -0.328d_b mined 0.423min_cos_consec (site) 0.7296min_cos_anchor (site) 0.7296dataset emolialang zhspeaker ZH_B00061_S00609total 36.1schain gain +1.0 dBseam step 1.5 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · balanced body, quiet background, slightly relaxed, fairly steady
(confusion, doubt · normal-paced, normally alert, some disfluency, casual)那这种情况下的话,只要是债权人或者说知情申请人完成了这么一个。
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as confusion, doubt; style: casual, conversational; below-average recording, quiet background; genuineness 4.0/6; vocal-burst blend 3.4/10; 6.3s, ZH.
ZH_B00061_S00609_W000078 · in -13.9 dBFS · gain -6.1 dB · emolia-03891
(doubt ·measured, normally alert, some disfluency, monologue)这个这个这提供这个抽逃出资行为的初步证据就行,或者我怀疑你抽逃出事了,这个出过证据东西。
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as doubt; style: monologue, storytelling; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 1.2/10; 10.0s, ZH.
ZH_B00061_S00609_W000079 · in -16.2 dBFS · gain -3.8 dB · emolia-03891
(disappointment, intoxication altered states of consciousness· measured, subdued, frequent disfluency, monologue)就唐山市赤城商贸公司北京二十一世纪营饮有限公司执行异议案,最高法院表明的证明责任的分配就是这么一个原则。
full caption & clip details
An elderly masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as disappointment, intoxication altered states of consciousness; style: monologue, whispered; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 2.9/10; 11.6s, ZH.
ZH_B00061_S00609_W000080 · in -15.7 dBFS · gain -4.3 dB · emolia-03891
(malevolence malice· measured, energised, some disfluency, authoritative)那作为被申请人被被告的公这个公司股东怎么来证明自己没有抽到厨资呢?比方说我用验资报告行不行?
full caption & clip details
A middle-aged masculine voice; delivery is energised, measured, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, rough, balanced body; clear, some disfluency, moderate pitch range, normal breath; affect is neutral, slightly dominant, fairly guarded; reads as malevolence malice; style: authoritative, monologue; average recording, quiet background; genuineness 2.7/6; vocal-burst blend 2.6/10; 8.7s, ZH.
ZH_B00061_S00609_W000081 · in -14.1 dBFS · gain -5.9 dB · emolia-03891
Fear ↓ / Sexual Lust ↑identity +0.09emotion 41 % proxy_spearman__PXR__T0.25__C0.25__INTERNAL · #12
This chain comes from the proxy rule: the same two-sided test as above, but because Sexual Lust is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Sexual Lust clearly present — 0.69, higher than 69 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.28.
At the same time Fear goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.69 (higher than 69 % of clips in this corpus), a change of -0.30. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.17, then +0.11 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.85 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.86 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.85. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 39 s · en · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.691 before conversion and 0.783 after — it rose by 0.092. Neighbour-to-neighbour the worst pair went 0.751 → 0.783. (The earlier render, with segment 1 left raw, scores 0.656 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sexual Lust moved +0.280 in the original and +0.116 after conversion — 41 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Fear, -0.298 became -0.314.
Quality. Mean predicted overall quality across the segments went 3.06 → 3.22 (+0.16) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.691 → 0.783+0.092identity cos neighbours 0.751 → 0.783d_b rescored +0.280 → +0.116d_a rescored -0.298 → -0.314d_a mined -0.298d_b mined 0.280min_cos_consec (site) 0.8608min_cos_anchor (site) 0.8525dataset emolialang enspeaker EN_B00001_S07245total 38.0schain gain +2.6 dBseam step 1.7 dBcrossfades 150/100 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, quiet background, neutral tension, moderately variable
(fear, interest, triumph · normal-paced, energised, normal breath, casual)Security violation shut down. I'm going to do show port security and I can specify, let me see the interface. I'll do interface GI four zero 39 and get some more info on that. There we go. Security enabled current port status. It's secure shut down. That hacker has been thwarted. He's not getting a thing from us. Woo. I'm pulling my Raspberry Pi back in. Now, one thing you'll notice is that
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is mildly positive, slightly dominant, slightly guarded; reads as fear, interest, triumph; style: casual, conversational; good recording, quiet background; mildly explicit content; genuineness 3.7/6; vocal-burst blend 5.4/10; 22.2s, EN.
EN_B00001_S07245_W000082 · in -23.7 dBFS · gain +3.7 dB · emolia-00285
(teasing, confusion, helplessness·brisk, energised, light breath, conversational)It doesn't come back up. How do you fix that once a violation's happened? And also if I do show IP interface brief and I only want to see my...
full caption & clip details
An adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as teasing, confusion, helplessness; style: conversational, storytelling; good recording, quiet background; genuineness 3.0/6; vocal-burst blend 4.1/10; 7.5s, EN.
EN_B00001_S07245_W000083 · in -25.5 dBFS · gain +5.5 dB · emolia-00285
(sexual lust, intoxication altered states of consciousness, amusement·normal-paced, normally alert, light breath, casual)One port there. I'll include GI four zero three nine. Oh wait, just four zero three nine. Sorry. Yeah. It is down down. You can also do show interfaces.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as sexual lust, intoxication altered states of consciousness, amusement; style: casual, conversational; good recording, quiet background; genuineness 4.6/6; vocal-burst blend 4.5/10; 8.7s, EN.
EN_B00001_S07245_W000084 · in -28.4 dBFS · gain +8.4 dB · emolia-00285
This chain comes from the proxy rule: the same two-sided test as above, but because Thankfulness Gratitude is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Thankfulness Gratitude clearly present — 0.64, higher than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.35.
At the same time Doubt goes the other way, from 0.93 (higher than 93 % of clips in this corpus) to 0.60 (higher than 60 % of clips in this corpus), a change of -0.33. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.23, then +0.11 — a fairly even climb, though some clips carry more of the change than others.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.82 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.83 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.82. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 29 s · en · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.783 before conversion and 0.775 after — it fell by 0.008. Neighbour-to-neighbour the worst pair went 0.783 → 0.775. (The earlier render, with segment 1 left raw, scores 0.607 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Thankfulness Gratitude moved +0.345 in the original and +0.250 after conversion — 72 % of the delta retained, which is most of it. On the other named axis, Doubt, -0.326 became -0.559.
Quality. Mean predicted overall quality across the segments went 2.74 → 3.17 (+0.43) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.783 → 0.775-0.008identity cos neighbours 0.783 → 0.775d_b rescored +0.345 → +0.250d_a rescored -0.326 → -0.559d_a mined -0.326d_b mined 0.345min_cos_consec (site) 0.8326min_cos_anchor (site) 0.8167dataset emolialang enspeaker EN_Dk94f3-5wVAtotal 28.4schain gain +4.1 dBseam step 0.7 dBcrossfades 150/150 ms
Script — 3 chunks, 3 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, normal-paced, normally alert, average clarity
(doubt · slightly relaxed, fairly steady, frequent disfluency, casual)(low mumble) Uh, whether it's even class, right? I mean, I'm getting used to this new format. I've been...
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as doubt; style: casual, conversational; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 2.3/10; 5.7s, EN.
EN_Dk94f3-5wVA_W000010 · in -18.8 dBFS · gain -1.2 dB · emolia-01596
(amusement, pleasure ecstasy, intoxication altered states of consciousness·neutral tension, moderately variable, some disfluency, casual)I've been joking with my wife, we've been watching some of the comedians that are on now that are broadcasting from their homes. And I'm like, Trevor Noah's not that funny at home either, right? So there is something lost when I'm sort of (low mumble) trying to teach into this little green light.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as amusement, pleasure ecstasy, intoxication altered states of consciousness; style: casual, monologue; good recording, quiet background; genuineness 4.0/6; vocal-burst blend 6.9/10; 13.3s, EN.
EN_Dk94f3-5wVA_W000011 · in -20.7 dBFS · gain +0.7 dB · emolia-01596
(thankfulness gratitude, contentment, pleasure ecstasy · neutral tension, fairly steady, some disfluency, conversational)(low mumble) Uh, that's sitting at the top of my monitor. (low mumble) Uhm, I miss you guys. And you know, the chance to be with you is something that's really fun for me. (low mumble) Uhm, and I know that the course staff enjoys that.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as thankfulness gratitude, contentment, pleasure ecstasy; style: conversational, casual; good recording, quiet background; genuineness 3.9/6; vocal-burst blend 5.8/10; 9.9s, EN.
EN_Dk94f3-5wVA_W000012 · in -20.5 dBFS · gain +0.5 dB · emolia-01596
This chain comes from the proxy rule: the same two-sided test as above, but because Confusion is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Confusion clearly present — 0.73, higher than 73 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.27.
At the same time Contempt goes the other way, from 0.91 (higher than 91 % of clips in this corpus) to 0.53 (higher than 53 % of clips in this corpus), a change of -0.38. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.14, then +0.13 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.75 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.75 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.75, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 20 s · zh · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.772 before conversion and 0.709 after — it fell by 0.063. Neighbour-to-neighbour the worst pair went 0.640 → 0.735. (The earlier render, with segment 1 left raw, scores 0.761 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Confusion moved +0.267 in the original and +0.790 after conversion — 296 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Contempt, -0.407 became -0.374.
Quality. Mean predicted overall quality across the segments went 2.73 → 2.95 (+0.22) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.772 → 0.709-0.063identity cos neighbours 0.640 → 0.735d_b rescored +0.267 → +0.790d_a rescored -0.407 → -0.374d_a mined -0.377d_b mined 0.267min_cos_consec (site) 0.7505min_cos_anchor (site) 0.7505dataset emolialang zhspeaker ZH_B00018_S08581total 19.7schain gain +2.1 dBseam step 0.8 dBcrossfades 150/100 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · fairly smooth, average recording, clear
This chain comes from the proxy rule: the same two-sided test as above, but because Concentration is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Concentration clearly present — 0.64, higher than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.33.
At the same time Shame goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.66 (higher than 66 % of clips in this corpus), a change of -0.32. Both halves had to happen for this chain to qualify.
It takes 5 clips to get there. Clip to clip the moves are -0.05, then +0.13, then +0.12, then +0.14 — not a clean run: step 1 moves back the other way by 0.05 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.30 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.32 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.30, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
5 clips · 73 s · no · eurospeech
What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.288 before conversion and 0.828 after — it rose by 0.539. Neighbour-to-neighbour the worst pair went 0.302 → 0.767. (The earlier render, with segment 1 left raw, scores 0.680 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.330 in the original and +0.489 after conversion — 148 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Shame, -0.324 became -0.159.
Quality. Mean predicted overall quality across the segments went 3.02 → 3.34 (+0.32) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.288 → 0.828+0.539identity cos neighbours 0.302 → 0.767d_b rescored +0.330 → +0.489d_a rescored -0.324 → -0.159d_a mined -0.324d_b mined 0.331min_cos_consec (site) 0.3156min_cos_anchor (site) 0.3032dataset eurospeechlang nospeaker norway_10591-2total 71.3schain gain +0.9 dBseam step 2.2 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · average recording, quiet background
(shame, bitterness, disappointment · brisk, energised, neutral tension, ranting)Jeg tenker at vi skal ha flere debatter om barnevern i året som kommer, og som komité vet vi at akkurat disse sakene knyttet til barnevern er de viktigste sakene vi som komité behandler.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, balanced body; clear, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, fairly guarded; reads as shame, bitterness, disappointment; style: ranting, cartoonish; average recording, quiet background; genuineness 2.1/6; vocal-burst blend 3.0/10; 13.3s, NO.
norway_10591-2_10567056_10580400 · in -23.7 dBFS · gain +3.7 dB · eurospeech-02230
(infatuation, bitterness, impatience and irritability· brisk, energised, neutral tension, storytelling)vi kan ha ulike syn på prioriteringer og hvilke grep som er best å ta, men vi trenger virkelig ikke å så tvil om hverandres engasje ment eller intensjoner.
full caption & clip details
A child feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; clear, some disfluency, very wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as infatuation, bitterness, impatience and irritability; style: storytelling, dramatic; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 2.5/10; 12.3s, NO.
norway_10591-2_10580400_10592704 · in -22.7 dBFS · gain +2.7 dB · eurospeech-02230
(shame, sexual lust, affection· brisk, energised, neutral tension, ranting)en gang for alle at alle representanter her, uavhengig av parti – og ikke minst barne- og familieministeren selv – er genuint opptatt av at vi må gjøre alt vi kan for å styrke og for bedre barnevernet.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, balanced body; clear, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as shame, sexual lust, affection; style: ranting, dramatic; average recording, quiet background; genuineness 2.0/6; vocal-burst blend 4.3/10; 13.3s, NO.
norway_10591-2_10592704_10606047 · in -23.3 dBFS · gain +3.3 dB · eurospeech-02230
(normal-paced, normally alert, slightly relaxed, casual)jeg kan svare (yawn) representanten Lossius med at jeg er ikke alene. Det er ikke representanten Kari Hen riksen som kommer med de utfallene mot regjeringa. Barneombudet har sagt under en budsjetthøring her i Stortinget at det er en styringskrise i barnevernet.
full caption & clip details
A middle-aged somewhat feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 1.4/10; 17.5s, NO.
norway_10591-2_10637616_10655088 · in -27.6 dBFS · gain +7.6 dB · eurospeech-02230
(concentration, confusion·measured, very low-energy, neutral tension, storytelling)Barnevernsbarna og Barneombudet har uttrykt alvorlig bekymring for rettssikkerheten i barnevernet. Den be kymringen deler jeg. Gang på gang har opposisjonen bedt
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, measured, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, balanced body; average clarity, frequent disfluency, wide pitch range, audible breath; affect is mildly positive, slightly dominant, slightly guarded; reads as concentration, confusion; style: storytelling, monologue; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 1.5/10; 15.7s, NO.
norway_10591-2_10655088_10670752 · in -25.5 dBFS · gain +5.5 dB · eurospeech-02230
This chain comes from the proxy rule: the same two-sided test as above, but because Relief is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Relief around average — 0.52, higher than 52 % of clips in this corpus — and ends with it at the very top of the corpus at 0.92, higher than 92 % of clips in this corpus. That is a total rise of 0.40.
At the same time Fatigue Exhaustion goes the other way, from 0.77 (higher than 77 % of clips in this corpus) to 0.40 (lower than 60 % of clips in this corpus), a change of -0.37. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.20, then +0.20 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.83 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.83 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.83. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 19 s · zh · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.841 before conversion and 0.861 after — it rose by 0.020. Neighbour-to-neighbour the worst pair went 0.839 → 0.819. (The earlier render, with segment 1 left raw, scores 0.746 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Relief moved +0.401 in the original and +0.405 after conversion — 101 % of the delta retained, which is essentially all of it. On the other named axis, Fatigue Exhaustion, -0.369 became -0.394.
Quality. Mean predicted overall quality across the segments went 2.92 → 3.10 (+0.18) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.841 → 0.861+0.020identity cos neighbours 0.839 → 0.819d_b rescored +0.401 → +0.405d_a rescored -0.369 → -0.394d_a mined -0.369d_b mined 0.401min_cos_consec (site) 0.8316min_cos_anchor (site) 0.8316dataset emolialang zhspeaker ZH_B00002_S07157total 18.5schain gain +1.9 dBseam step 1.2 dBcrossfades 150/100 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, average recording, no background noise, normally alert, slightly relaxed, fairly steady
(measured, some disfluency, somewhat unclear, monologue)有的时候你在跟人交流的时候,看有人皱起眉头,或者呢脸上的表情不好。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, didactic; average recording, no background noise; genuineness 2.1/6; vocal-burst blend 1.2/10; 6.1s, ZH.
ZH_B00002_S07157_W000010 · in -19.3 dBFS · gain -0.7 dB · emolia-03299
(normal-paced, no disfluency, average clarity, formal)这些身体语言似乎都在告诉你,他们不喜欢你说的。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal; average recording, no background noise; genuineness 2.9/6; vocal-burst blend 1.9/10; 3.8s, ZH.
ZH_B00002_S07157_W000011 · in -20.3 dBFS · gain +0.3 dB · emolia-03299
(relief·measured, some disfluency, somewhat unclear, didactic)你就知道呢,他们可能要批评你,或者呢心里反对你,有的人就被这种身体语言给吓得趴下了。
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as relief; style: didactic, monologue; average recording, no background noise; genuineness 2.2/6; vocal-burst blend 1.1/10; 8.9s, ZH.
ZH_B00002_S07157_W000012 · in -17.3 dBFS · gain -2.7 dB · emolia-03299
This chain comes from the proxy rule: the same two-sided test as above, but because Hope Enthusiasm Optimism is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Hope Enthusiasm Optimism clearly present — 0.59, higher than 59 % of clips in this corpus — and ends with it strongly present at 0.85, higher than 85 % of clips in this corpus. That is a total rise of 0.26.
At the same time Interest goes the other way, from 0.91 (higher than 91 % of clips in this corpus) to 0.53 (higher than 53 % of clips in this corpus), a change of -0.38. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.13, then +0.13 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.66 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.81 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.66, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 45 s · en · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.697 before conversion and 0.709 after — it rose by 0.013. Neighbour-to-neighbour the worst pair went 0.800 → 0.799. (The earlier render, with segment 1 left raw, scores 0.679 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Hope Enthusiasm Optimism moved +0.260 in the original and +0.400 after conversion — 154 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Interest, -0.379 became -0.314.
Quality. Mean predicted overall quality across the segments went 2.92 → 3.12 (+0.19) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.697 → 0.709+0.013identity cos neighbours 0.800 → 0.799d_b rescored +0.260 → +0.400d_a rescored -0.379 → -0.314d_a mined -0.379d_b mined 0.260min_cos_consec (site) 0.8059min_cos_anchor (site) 0.6613dataset emolialang enspeaker EN_B00049_S03248total 44.3schain gain +4.7 dBseam step 1.3 dBcrossfades 100/100 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · fairly smooth, balanced body, good recording, slightly relaxed, clear
(interest · measured, subdued, steady, didactic)Number 16 is price. You can say a high price or an exorbitant price, which means an overly high price, a low price, a reasonable price, the right price, asking price, half price, retail price,
full caption & clip details
A young adult feminine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is slightly cool, slightly bright, fairly smooth, balanced body; clear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly positive, neutral stance, neutral openness; reads as interest; style: didactic, whispered; good recording, quiet background; genuineness 0.8/6; vocal-burst blend 0.6/10; 19.7s, EN.
EN_B00049_S03248_W000033 · in -22.2 dBFS · gain +2.2 dB · emolia-01202
(measured, normally alert, fairly steady, whispered)Cut prices, you can quote someone on a price, something can go up in price, and you can talk about a price range, a price increase, a price rise, a price cut, and a price war.
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; clear, frequent disfluency, wide pitch range, normal breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: whispered, didactic; good recording, quiet background; genuineness 1.0/6; vocal-burst blend 0.6/10; 14.8s, EN.
EN_B00049_S03248_W000034 · in -22.7 dBFS · gain +2.7 dB · emolia-01202
(normal-paced, normally alert, steady, dramatic)You can increase production, you can have production costs, you can talk about mass production, you can start production, and you can stop production.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: dramatic, formal; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 0.6/10; 10.1s, EN.
EN_B00049_S03248_W000035 · in -23.3 dBFS · gain +3.3 dB · emolia-01202
This chain comes from the proxy rule: the same two-sided test as above, but because Embarrassment is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Embarrassment clearly present — 0.70, higher than 70 % of clips in this corpus — and ends with it at the very top of the corpus at 0.96, higher than 96 % of clips in this corpus. That is a total rise of 0.26.
At the same time Astonishment Surprise goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.66 (higher than 66 % of clips in this corpus), a change of -0.30. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.03, then +0.23 — a slow start, with most of the change arriving in the final step.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.80 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.86 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.80. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 48 s · en · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.710 before conversion and 0.665 after — it fell by 0.044. Neighbour-to-neighbour the worst pair went 0.760 → 0.674. (The earlier render, with segment 1 left raw, scores 0.585 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Embarrassment moved +0.257 in the original and +0.247 after conversion — 96 % of the delta retained, which is essentially all of it. On the other named axis, Astonishment Surprise, -0.296 became -0.443.
Quality. Mean predicted overall quality across the segments went 2.89 → 3.18 (+0.28) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.710 → 0.665-0.044identity cos neighbours 0.760 → 0.674d_b rescored +0.257 → +0.247d_a rescored -0.296 → -0.443d_a mined -0.296d_b mined 0.257min_cos_consec (site) 0.8598min_cos_anchor (site) 0.8015dataset emolialang enspeaker EN_B00008_S00955total 47.0schain gain +4.5 dBseam step 1.1 dBcrossfades 150/150 ms
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, light breath
(astonishment surprise, intoxication altered states of consciousness, impatience and irritability · normal-paced, subdued, neutral tension, casual)The, the initial reaction from, from Leif and Seth was, make, put those guys, you know, that's two, you know, that's two other squads, you know, that's, we, we don't want them in our platoon. And I was like, (exhausted groan) uh, no. What we're gonna do is we're gonna take one or two of them and put them in each one of our fire teams and make them part of our unit. And, you know, they objected.
full caption & clip details
An adult masculine voice; delivery is subdued, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, slightly guarded; reads as astonishment surprise, intoxication altered states of consciousness, impatience and irritability; style: casual, conversational; good recording, quiet background; genuineness 5.1/6; vocal-burst blend 9.2/10; 24.3s, EN.
EN_B00008_S00955_W000200 · in -16.8 dBFS · gain -3.2 dB · emolia-00417
(elation, anger, interest·brisk, energised, neutral tension, casual)What happens? What happens is that those people get individually, it's very easy for a four or five man fire team to absorb one or two people. Yeaah. Hey cool, what's your name? Fred? Cool. Right on. You're good. Hey, (ahem) I'm Bill. Let me know. Here's where I'm at. Here's how we do our head count. Here's what we're looking for. Boom. Integration was almost immediate.
full caption & clip details
An adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, slightly guarded; reads as elation, anger, interest; style: casual, conversational; good recording, quiet background; genuineness 4.8/6; vocal-burst blend 10.0/10; 19.2s, EN.
EN_B00008_S00955_W000201 · in -15.7 dBFS · gain -4.3 dB · emolia-00417
(embarrassment, impatience and irritability·normal-paced, normally alert, slightly relaxed, casual)Cause they weren't really fire teams, they were a little bit bigger now, so I called them sections. And...
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as embarrassment, impatience and irritability; style: casual, conversational; good recording, no background noise; genuineness 3.4/6; vocal-burst blend 3.0/10; 3.9s, EN.
EN_B00008_S00955_W000202 · in -15.9 dBFS · gain -4.1 dB · emolia-00417
This chain comes from the proxy rule: the same two-sided test as above, but because Hope Enthusiasm Optimism is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Hope Enthusiasm Optimism clearly present — 0.67, higher than 67 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.32.
At the same time Doubt goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.49 (lower than 51 % of clips in this corpus), a change of -0.50. Both halves had to happen for this chain to qualify.
It takes 4 clips to get there. Clip to clip the moves are +0.21, then -0.11, then +0.21 — not a clean run: step 2 moves back the other way by 0.11 before the chain recovers.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.71 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.77 against each other.
Voice consistency: these clips are separate recordings joined together and the match is loose (0.71, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
4 clips · 83 s · en · podcast
What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.715 before conversion and 0.787 after — it rose by 0.073. Neighbour-to-neighbour the worst pair went 0.753 → 0.761. (The earlier render, with segment 1 left raw, scores 0.636 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Hope Enthusiasm Optimism moved +0.320 in the original and +0.421 after conversion — 131 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Doubt, -0.501 became -0.262.
Quality. Mean predicted overall quality across the segments went 2.74 → 3.08 (+0.33) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.715 → 0.787+0.073identity cos neighbours 0.753 → 0.761d_b rescored +0.320 → +0.421d_a rescored -0.501 → -0.262d_a mined -0.501d_b mined 0.320min_cos_consec (site) 0.7693min_cos_anchor (site) 0.7132dataset podcastlang enspeaker 977902total 82.4schain gain +3.8 dBseam step 1.2 dBcrossfades 150/150/150 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a child masculine voice · quiet background, average clarity
(doubt, bitterness, disappointment · normal-paced, normally alert, neutral tension, cartoonish)(ahem) Uh, they might get hypoglycemic, the sugar can go low because remember, I'm taking up 70% of your plasma, and that's where all your sugar is. So that's why we always make sure people eat before we do the procedure. And they might become volume depleted. Their volume may fall a little bit because of the exchange that we're doing, but we usually give them saline, and that doesn't really happen very often. But other than that, there's very little risk, and we've seen very few adverse reactions on the hundreds and hundreds and thousands that we've done in this office.
full caption & clip details
A child masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is slightly cool, dark, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as doubt, bitterness, disappointment; style: cartoonish, monologue; average recording, quiet background; genuineness 2.6/6; vocal-burst blend 4.6/10; 28.4s, EN.
977902_00266432 · in -29.3 dBFS · gain +9.3 dB · podcast-01155
(jealousy and envy, anger, bitterness ·measured, subdued, neutral tension, casual)They they they come off very very easily, and people are quite amazed. This is one of the few things that I've been a doctor long enough that I have people sending me testimonials without me asking them. They're like, you need to post this. And they can go to YouTube and our website and see all these testimonials that have with people, healthy people, people who are not so healthy.
full caption & clip details
An elderly masculine voice; delivery is subdued, measured, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, frequent disfluency, moderate pitch range, normal breath; affect is mildly negative, slightly dominant, slightly guarded; reads as jealousy and envy, anger, bitterness; style: casual, monologue; average recording, quiet background; genuineness 4.9/6; vocal-burst blend 6.5/10; 21.6s, EN.
977902_00269268 · in -29.2 dBFS · gain +9.2 dB · podcast-01154
(affection, thankfulness gratitude, pride· measured, normally alert, slightly relaxed, didactic)I'll tell you one of the things that I'm taking away from this on the longevity part towards CPE is everybody's improving with it, even the people who are very healthy, but the people who are
full caption & clip details
A child masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, thin; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as affection, thankfulness gratitude, pride; style: didactic, narration; good recording, quiet background; genuineness 1.7/6; vocal-burst blend 0.2/10; 9.7s, EN.
977902_00271424 · in -28.2 dBFS · gain +8.2 dB · podcast-01156
(hope enthusiasm optimism, affection, contentment·normal-paced, normally alert, neutral tension, casual)That's kind of what we always see, right? The people with the worst diets do the best when we clean them up. People have pretty good diet, they get a little bit better, but and so it's it's pretty much but the but the rate at which the improvements happen on across the board for people has been very satisfying. Very satisfying for me. That is fascinating. Wow, what an amazing explanation of what TPE is.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, slightly guarded; reads as hope enthusiasm optimism, affection, contentment; style: casual, conversational; average recording, quiet background; genuineness 4.9/6; vocal-burst blend 6.7/10; 23.2s, EN.
977902_00272824 · in -29.7 dBFS · gain +9.7 dB · podcast-01160
This chain comes from the proxy rule: the same two-sided test as above, but because Fear is not one of the emotions that ramps cleanly on its own, the per-step cap was applied to a stand-in (“proxy”) axis that tracks it.
The chain starts with Fear around average — 0.56, higher than 56 % of clips in this corpus — and ends with it strongly present at 0.90, higher than 90 % of clips in this corpus. That is a total rise of 0.34.
At the same time Emotional Numbness goes the other way, from 0.95 (higher than 95 % of clips in this corpus) to 0.54 (higher than 54 % of clips in this corpus), a change of -0.41. Both halves had to happen for this chain to qualify.
It takes 3 clips to get there. Clip to clip the moves are +0.13, then +0.21 — an even, steady climb — each clip carries about the same share.
No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.
Same speaker? The least similar clip scores 0.83 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.83 against each other.
Voice consistency: these clips are separate recordings joined together, matching at 0.83. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.
3 clips · 27 s · en · emolia
What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.
Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.742 before conversion and 0.732 after — it fell by 0.010. Neighbour-to-neighbour the worst pair went 0.742 → 0.732. (The earlier render, with segment 1 left raw, scores 0.622 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.
Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fear moved +0.339 in the original and +0.350 after conversion — 103 % of the delta retained, which is essentially all of it. On the other named axis, Emotional Numbness, -0.414 became -0.667.
Quality. Mean predicted overall quality across the segments went 2.14 → 2.96 (+0.82) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.
corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.742 → 0.732-0.010identity cos neighbours 0.742 → 0.732d_b rescored +0.339 → +0.350d_a rescored -0.414 → -0.667d_a mined -0.414d_b mined 0.338min_cos_consec (site) 0.8283min_cos_anchor (site) 0.8283dataset emolialang enspeaker EN_t9wobHtto6Etotal 26.3schain gain +1.4 dBseam step 4.3 dBcrossfades 150/150 ms
Script — 3 chunks, 3 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normal-paced, normally alert, slightly relaxed, fairly steady
(emotional numbness, helplessness · some disfluency, average clarity, monologue)I'm just sort of a top line. (ahem) I'm supportive of the Planning Commission recommendations. (ahem)
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, helplessness; style: monologue; good recording, no background noise; genuineness 3.0/6; vocal-burst blend 1.3/10; 5.0s, EN.
EN_t9wobHtto6E_W000992 · in -16.5 dBFS · gain -3.5 dB · emolia-01936
(contentment· some disfluency, average clarity, monologue)Obviously reviewed not just all their materials, but listened to (low mumble) a lot of their deliberations and I thought they were very thoughtful in the process and I thought the recommendations and changes they made (ahem) were the right ones.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contentment; style: monologue; average recording, no background noise; genuineness 3.3/6; vocal-burst blend 2.1/10; 10.0s, EN.
EN_t9wobHtto6E_W000993 · in -17.1 dBFS · gain -2.9 dB · emolia-01936
(frequent disfluency, somewhat unclear, monologue)I do support some of the (low mumble) clarifications that are coming forward from (ahem) county staff, not necessarily all of them. When we get down to it, (ahem) uhm, we can get into those weeds. But I just think from a general, (low mumble) uhm,
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 3.2/10; 11.7s, EN.
EN_t9wobHtto6E_W000994 · in -17.3 dBFS · gain -2.7 dB · emolia-01936