k-B1-k3 — voice-corrected

B1 at chain length k=3, all corpora, at the mining floor.

This is the voice-conversion-corrected version of this tier. The original, uncorrected page is still there and unchanged: t_k-B1-k3.html. Every card below carries both renders so you can switch between them without leaving the page. What was done, and what it measured →
20chains converted
40segments re-voiced
0.795 → 0.811median worst-to-anchor identity cosine
71 %median emotion delta retained
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken.

(underlined, plain · delivery, style) — the tag before the words. Emotions first, then how it is delivered. Underlined descriptors are the ones that change across this chain — anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words are a real non-speech sound, printed where it happens.

The full generated caption for any clip is under “full caption & clip details”. Its perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.

The tags and captions describe the ORIGINAL clips — they are the corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the converted audio are the ones printed in each card's own paragraph and chips.
Sexual Lust(unconstrained axis: Pain)identity −0.05 emotion 208 %   k-B1-k3 · #1

This chain comes from the one-sided rule: only Sexual Lust had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Sexual Lust around average — 0.54, higher than 54 % of clips in this corpus — and ends with it at the very top of the corpus at 0.92, higher than 92 % of clips in this corpus. That is a total rise of 0.39.

Nothing was asked of the other axis, and in fact Pain barely moves at all, sitting near 0.90 throughout.

It takes 3 clips to get there. Clip to clip the moves are +0.16, then +0.23 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 14 s · en · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.709 before conversion and 0.658 after — it fell by 0.051. Neighbour-to-neighbour the worst pair went 0.755 → 0.697. (The earlier render, with segment 1 left raw, scores 0.538 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sexual Lust moved +0.385 in the original and +0.802 after conversion — 208 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Pain, +0.013 became +0.011.

Quality. Mean predicted overall quality across the segments went 2.46 → 2.72 (+0.26) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.709 → 0.658 -0.051identity cos neighbours 0.755 → 0.697d_b rescored +0.385 → +0.802d_a rescored +0.013 → +0.011d_a mined 0.013d_b mined 0.385min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_J0oTYbacO4Etotal 13.6schain gain +2.1 dBseam step 2.3 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, slightly bright, fairly smooth, balanced body, good recording, normal-paced, normally alert, slightly relaxed
(casual, monologue) And the hypothalamus arouses the autonomic nervous system. Your body begins to reserve fluids.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: casual, monologue; good recording, quiet background; genuineness 2.1/6; vocal-burst blend 0.6/10; 6.8s, EN.
EN_J0oTYbacO4E_W000072 · in -18.5 dBFS · gain -1.5 dB · emolia-00574
(disgust, emotional numbness · monologue, casual) And controls the release of saliva, tears, and gastric acid.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as disgust, emotional numbness; style: monologue, casual; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 2.3/10; 3.8s, EN.
EN_J0oTYbacO4E_W000073 · in -17.4 dBFS · gain -2.6 dB · emolia-00574
(sexual lust, pain · casual, monologue) It creates cortisol, which actually assists your blood and clotting.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as sexual lust, pain; style: casual, monologue; good recording, no background noise; genuineness 2.2/6; vocal-burst blend 2.6/10; 3.5s, EN.
EN_J0oTYbacO4E_W000074 · in -18.8 dBFS · gain -1.2 dB · emolia-00574
Contempt(unconstrained axis: Relief)identity +0.03 emotion 43 %   k-B1-k3 · #2

This chain comes from the one-sided rule: only Contempt had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Contempt clearly present — 0.70, higher than 70 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.25.

Nothing was asked of the other axis, and in fact Relief drifts down from 1.00 to 0.53 (-0.46), which the rule did not require.

It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.01 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.73 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.73 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.73, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 47 s · en · podcast

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.732 before conversion and 0.764 after — it rose by 0.033. Neighbour-to-neighbour the worst pair went 0.732 → 0.822. (The earlier render, with segment 1 left raw, scores 0.612 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contempt moved +0.255 in the original and +0.109 after conversion — 43 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Relief, -0.432 became -0.318.

Quality. Mean predicted overall quality across the segments went 2.56 → 2.96 (+0.41) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.732 → 0.764 +0.033identity cos neighbours 0.732 → 0.822d_b rescored +0.255 → +0.109d_a rescored -0.432 → -0.318d_a mined -0.462d_b mined 0.254min_cos_consec (site) 0.7318min_cos_anchor (site) 0.7318dataset podcastlang enspeaker 670383total 46.8schain gain +5.6 dBseam step 0.9 dBcrossfades 150/100 ms
Script — 3 chunks, 3 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · slightly bright, quiet background, energised, neutral tension, moderately variable, wide pitch range, normal breath
(relief, astonishment surprise, triumph · normal-paced, some disfluency, average clarity, casual) we (surprised gasp) uh we understood that Shep was good because people were like when he came back, people were like, (surprised gasp) Oh, is that Shep? Like, it was like stuff like that. But with Flip.
full caption & clip details
A young adult feminine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as relief, astonishment surprise, triumph; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 4.9/6; vocal-burst blend 5.3/10; 11.0s, EN.
670383_00326648 · in -18.5 dBFS · gain -1.5 dB · podcast-00991
(affection, thankfulness gratitude, infatuation · brisk, some disfluency, average clarity, conversational) connected well. And I think in your defense, not coming down on you, (ahem) uh, but I feel like you love this movie so much and you've seen it so much to where and you know enough about basketball to
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as affection, thankfulness gratitude, infatuation; style: conversational, casual; average recording, quiet background; genuineness 4.6/6; vocal-burst blend 6.0/10; 12.2s, EN.
670383_00330488 · in -18.2 dBFS · gain -1.8 dB · podcast-03183
(contempt, amusement, sourness · normal-paced, frequent disfluency, somewhat unclear, casual) does not know basketball, who's only seen who's seen this for like the first time, it was just like I'm very confused on who is who and what is what and why is why so you know but let's let's go to (low mumble) um let's talk about (low mumble) um the tournament
full caption & clip details
A young adult feminine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, slightly rough, slightly thin; somewhat unclear, frequent disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as contempt, amusement, sourness; style: casual, conversational; below-average recording, quiet background; mildly explicit content; genuineness 4.6/6; vocal-burst blend 4.3/10; 23.9s, EN.
670383_00332176 · in -17.9 dBFS · gain -2.1 dB · podcast-00283
Fatigue Exhaustion(unconstrained axis: Contemplation)identity −0.04 emotion 39 %   k-B1-k3 · #3

This chain comes from the one-sided rule: only Fatigue Exhaustion had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Fatigue Exhaustion around average — 0.57, higher than 57 % of clips in this corpus — and ends with it strongly present at 0.87, higher than 87 % of clips in this corpus. That is a total rise of 0.30.

Nothing was asked of the other axis, and in fact Contemplation drifts down from 0.75 to 0.48 (-0.27), which the rule did not require.

It takes 3 clips to get there. Clip to clip the moves are +0.14, then +0.16 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 20 s · zh · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.807 before conversion and 0.764 after — it fell by 0.043. Neighbour-to-neighbour the worst pair went 0.869 → 0.817. (The earlier render, with segment 1 left raw, scores 0.788 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fatigue Exhaustion moved +0.303 in the original and +0.118 after conversion — 39 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Contemplation, -0.272 became -0.094.

Quality. Mean predicted overall quality across the segments went 3.02 → 3.11 (+0.09) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.807 → 0.764 -0.043identity cos neighbours 0.869 → 0.817d_b rescored +0.303 → +0.118d_a rescored -0.272 → -0.094d_a mined -0.270d_b mined 0.303min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00017_S08426total 18.9schain gain +1.9 dBseam step 0.5 dBcrossfades 150/100 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, no background noise, normally alert, slightly relaxed, no disfluency
(normal-paced, fairly steady, moderate pitch range, formal) 因为我认为树象征着创业者,要专注,要扎根。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, narration; good recording, no background noise; genuineness 2.2/6; vocal-burst blend 2.8/10; 3.6s, ZH.
ZH_B00017_S08426_W000063 · in -20.3 dBFS · gain +0.3 dB · emolia-03447
(measured, fairly steady, moderate pitch range, narration) 在生命健康领域的创业更是如此,必须深深的扎根,才可能开花结果。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, formal; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 4.7/10; 6.0s, ZH.
ZH_B00017_S08426_W000064 · in -19.9 dBFS · gain -0.1 dB · emolia-03447
(measured, steady, fairly narrow pitch, monologue) 我自己见证过生命健康项目的创业过程。知道一个项目,从创意到产生新药或者医疗器械,要走多么漫长而艰难的道路。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; average recording, no background noise; genuineness 0.3/6; vocal-burst blend 3.0/10; 9.7s, ZH.
ZH_B00017_S08426_W000065 · in -19.8 dBFS · gain -0.2 dB · emolia-03447
Infatuation(unconstrained axis: Sexual Lust)identity −0.01 emotion 22 %   k-B1-k3 · #4

This chain comes from the one-sided rule: only Infatuation had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Infatuation around average — 0.51, right about the corpus median — and ends with it strongly present at 0.78, higher than 78 % of clips in this corpus. That is a total rise of 0.28.

Nothing was asked of the other axis, and in fact Sexual Lust drifts down from 0.82 to 0.01 (-0.81), which the rule did not require.

It takes 3 clips to get there. Clip to clip the moves are +0.06, then +0.22 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.91 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.91 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 14 s · zh · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.862 before conversion and 0.855 after — it fell by 0.007. Neighbour-to-neighbour the worst pair went 0.901 → 0.859. (The earlier render, with segment 1 left raw, scores 0.763 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Infatuation moved +0.277 in the original and +0.062 after conversion — 22 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Sexual Lust, -0.809 became -0.441.

Quality. Mean predicted overall quality across the segments went 2.67 → 2.97 (+0.29) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.862 → 0.855 -0.007identity cos neighbours 0.901 → 0.859d_b rescored +0.277 → +0.062d_a rescored -0.809 → -0.441d_a mined -0.809d_b mined 0.277min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00022_S07121total 13.3schain gain -0.1 dBseam step 1.0 dBcrossfades 150/100 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, fairly smooth, balanced body, no background noise, normally alert, slightly relaxed, no disfluency, moderate pitch range
(fast, fairly steady, slurred, formal) 追赶超越,是大西安当下最大的事业。
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; slurred, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, casual; average recording, no background noise; genuineness 1.2/6; vocal-burst blend 2.7/10; 3.2s, ZH.
ZH_B00022_S07121_W000013 · in -17.1 dBFS · gain -2.9 dB · emolia-03493
(normal-paced, steady, clear, formal) 需要咬定青山驰而不息的奋斗到底。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 2.4/10; 3.3s, ZH.
ZH_B00022_S07121_W000014 · in -17.7 dBFS · gain -2.3 dB · emolia-03493
(measured, fairly steady, clear, monologue) 我们要以更加谦虚的态度,更加过人的干劲,积极听取来自各方面的意见建议。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; average recording, no background noise; genuineness 0.6/6; vocal-burst blend 2.3/10; 7.2s, ZH.
ZH_B00022_S07121_W000015 · in -17.3 dBFS · gain -2.7 dB · emolia-03493
Longing(unconstrained axis: Thankfulness Gratitude)identity −0.01 emotion 73 %   k-B1-k3 · #5

This chain comes from the one-sided rule: only Longing had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Longing clearly present — 0.61, higher than 61 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.36.

Nothing was asked of the other axis, and in fact Thankfulness Gratitude drifts down from 0.76 to 0.24 (-0.52), which the rule did not require.

It takes 3 clips to get there. Clip to clip the moves are +0.11, then +0.24 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 27 s · zh · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.851 before conversion and 0.840 after — it fell by 0.011. Neighbour-to-neighbour the worst pair went 0.851 → 0.859. (The earlier render, with segment 1 left raw, scores 0.789 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Longing moved +0.356 in the original and +0.258 after conversion — 73 % of the delta retained, which is most of it. On the other named axis, Thankfulness Gratitude, -0.522 became -0.660.

Quality. Mean predicted overall quality across the segments went 3.17 → 3.27 (+0.10) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.851 → 0.840 -0.011identity cos neighbours 0.851 → 0.859d_b rescored +0.356 → +0.258d_a rescored -0.522 → -0.660d_a mined -0.522d_b mined 0.356min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00052_S04527total 25.9schain gain +0.5 dBseam step 0.8 dBcrossfades 150/100 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, no background noise, measured, normally alert, slightly relaxed
(steady, no disfluency, monologue, formal) 虚发洁白的老人,和蔼的笑了笑,和自己的老伴披上外衣,径直走了出去。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 1.5/10; 5.8s, ZH.
ZH_B00052_S04527_W000042 · in -14.9 dBFS · gain -5.1 dB · emolia-03799
(pain · fairly steady, no disfluency, monologue, narration) 就在两个老人就要打开木门的时候,李雪梅轻轻放下怀中的小婴儿身影闪动,和金砖一起在远方,静静的看着刺啦。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pain; style: monologue, narration; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 3.8/10; 9.8s, ZH.
ZH_B00052_S04527_W000043 · in -14.4 dBFS · gain -5.5 dB · emolia-03799
(longing, contemplation, sadness · fairly steady, some disfluency, whispered, monologue) (ahem) 木门被打开了,两个老人,却并没有看到什么人,只看见一个还在襁褓之中的婴儿,正放在自己家的门前,嗯,是一个婴儿。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as longing, contemplation, sadness; style: whispered, monologue; average recording, no background noise; genuineness 1.7/6; vocal-burst blend 4.2/10; 10.6s, ZH.
ZH_B00052_S04527_W000044 · in -14.8 dBFS · gain -5.2 dB · emolia-03799
Pain(unconstrained axis: Astonishment Surprise)identity −0.00 emotion 134 %   k-B1-k3 · #6

This chain comes from the one-sided rule: only Pain had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Pain around average — 0.54, higher than 54 % of clips in this corpus — and ends with it strongly present at 0.82, higher than 82 % of clips in this corpus. That is a total rise of 0.28.

Nothing was asked of the other axis, and in fact Astonishment Surprise drifts down from 0.89 to 0.52 (-0.37), which the rule did not require.

It takes 3 clips to get there. Clip to clip the moves are +0.17, then +0.10 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 28 s · en · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.818 before conversion and 0.816 after — it fell by 0.002. Neighbour-to-neighbour the worst pair went 0.799 → 0.816. (The earlier render, with segment 1 left raw, scores 0.699 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Pain moved +0.276 in the original and +0.369 after conversion — 134 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Astonishment Surprise, -0.368 became -0.325.

Quality. Mean predicted overall quality across the segments went 2.89 → 3.01 (+0.12) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.818 → 0.816 -0.002identity cos neighbours 0.799 → 0.816d_b rescored +0.276 → +0.369d_a rescored -0.368 → -0.325d_a mined -0.368d_b mined 0.276min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_B00010_S03415total 27.0schain gain +1.8 dBseam step 3.3 dBcrossfades 100/100 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, fairly smooth, balanced body, good recording, brisk, normally alert, some disfluency, light breath
(slightly relaxed, moderately variable, average clarity, casual) So Ashton has put about $10,000 in this year. (ahem) He does not currently receive social security, but we keep his assets low just in case for the future. Since he has income, it seems like he could open a Roth IRA.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; good recording, no background noise; genuineness 2.3/6; vocal-burst blend 3.4/10; 12.9s, EN.
EN_B00010_S03415_W000018 · in -19.0 dBFS · gain -1.0 dB · emolia-00455
(infatuation · slightly relaxed, fairly steady, average clarity, conversational) Or 3B. However, a Roth IRA is not an employer sponsored plan, so it would be fine to have both.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as infatuation; style: conversational, casual; good recording, quiet background; genuineness 3.3/6; vocal-burst blend 2.8/10; 5.8s, EN.
EN_B00010_S03415_W000019 · in -18.7 dBFS · gain -1.3 dB · emolia-00455
(neutral tension, fairly steady, clear, casual) This one's from Pamela in Georgia. What is a good alternative to Microsoft Office 365 for a work from home person needing a product like this for typing?
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; good recording, no background noise; genuineness 1.4/6; vocal-burst blend 2.4/10; 8.7s, EN.
EN_B00010_S03415_W000020 · in -18.3 dBFS · gain -1.7 dB · emolia-00455
Pride(unconstrained axis: Disappointment)identity +0.00 emotion 86 %   k-B1-k3 · #7

This chain comes from the one-sided rule: only Pride had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Pride clearly present — 0.69, higher than 69 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.30.

Nothing was asked of the other axis, and in fact Disappointment drifts down from 0.98 to 0.39 (-0.59), which the rule did not require.

It takes 3 clips to get there. Clip to clip the moves are +0.18, then +0.12 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 53 s · zh · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.920 before conversion and 0.925 after — it rose by 0.005. Neighbour-to-neighbour the worst pair went 0.890 → 0.923. (The earlier render, with segment 1 left raw, scores 0.876 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Pride moved +0.297 in the original and +0.256 after conversion — 86 % of the delta retained, which is most of it. On the other named axis, Disappointment, -0.589 became -0.595.

Quality. Mean predicted overall quality across the segments went 3.01 → 3.32 (+0.31) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.920 → 0.925 +0.005identity cos neighbours 0.890 → 0.923d_b rescored +0.297 → +0.256d_a rescored -0.589 → -0.595d_a mined -0.589d_b mined 0.298min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00077_S00711total 52.5schain gain +0.6 dBseam step 1.8 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · fairly smooth, thin, average recording, quiet background, normally alert, slightly relaxed, some disfluency, light breath
(disappointment, shame, interest · normal-paced, fairly steady, clear, monologue) 今天我们要聊的这部最后的决斗呢,它是根据美国中世纪文学教授,一个叫艾瑞克雅阁。他在零四年写的同名历史小说改编而来的那他的真实的历史原型确实是一三八六年十二月二十九日,这个诺曼的骑士叫round士。
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as disappointment, shame, interest; style: monologue, narration; average recording, quiet background; genuineness 1.6/6; vocal-burst blend 4.7/10; 18.9s, ZH.
ZH_B00077_S00711_W000002 · in -13.1 dBFS · gain -6.9 dB · emolia-04043
(disappointment, disgust, jealousy and envy · fast, moderately variable, average clarity, cartoonish) Around doga护士,以妻子玛格丽特德ga护士,也是由我非常非常喜欢的女演员朱迪科莫饰演的那这个妻子呢,玛格丽特其实遭到了一个国王护卫叫雅克勒格里斯zaclerk z也是由十六姐非常喜欢的亚当德莱福饰演的这个角色。
full caption & clip details
A child feminine voice; delivery is normally alert, fast, slightly relaxed, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as disappointment, disgust, jealousy and envy; style: cartoonish, storytelling; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 7.2/10; 19.4s, ZH.
ZH_B00077_S00711_W000003 · in -11.8 dBFS · gain -8.2 dB · emolia-04043
(pride, malevolence malice, hope enthusiasm optimism · brisk, moderately variable, clear, dramatic) 嗯,反正就是都是男神女神集结吧,就相当于是这个妻子遭到这个护卫侵犯之后,丈夫去申诉,最后审讯无果,那两个人就进行了法国历史上记载的最后一场就是比武审判。
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, thin; clear, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as pride, malevolence malice, hope enthusiasm optimism; style: dramatic, monologue; average recording, quiet background; genuineness 2.6/6; vocal-burst blend 5.7/10; 14.6s, ZH.
ZH_B00077_S00711_W000004 · in -11.6 dBFS · gain -8.4 dB · emolia-04043
Emotional Numbness(unconstrained axis: Infatuation)identity +0.05 emotion REVERSED   k-B1-k3 · #8

This chain comes from the one-sided rule: only Emotional Numbness had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Emotional Numbness around average — 0.56, higher than 56 % of clips in this corpus — and ends with it strongly present at 0.84, higher than 84 % of clips in this corpus. That is a total rise of 0.28.

Nothing was asked of the other axis, and in fact Infatuation drifts down from 0.87 to 0.65 (-0.22), which the rule did not require.

It takes 3 clips to get there. Clip to clip the moves are +0.10, then +0.17 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.88 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.92 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.88. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 29 s · sv · podcast

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.863 before conversion and 0.912 after — it rose by 0.049. Neighbour-to-neighbour the worst pair went 0.902 → 0.923. (The earlier render, with segment 1 left raw, scores 0.862 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

The emotional move did not survive. Re-scored end to end, Emotional Numbness moved +0.291 in the original and -0.194 after conversion — it changed direction. On this chain the corrected audio is not an improvement. On the other named axis, Infatuation, -0.221 became -0.201.

Quality. Mean predicted overall quality across the segments went 3.08 → 3.36 (+0.28) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.863 → 0.912 +0.049identity cos neighbours 0.902 → 0.923d_b rescored +0.291 → -0.194d_a rescored -0.221 → -0.201d_a mined -0.221d_b mined 0.278min_cos_consec (site) 0.9199min_cos_anchor (site) 0.8813dataset podcastlang svspeaker 855522total 28.7schain gain +0.2 dBseam step 1.3 dBcrossfades 150/100 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a child masculine voice · neutral-toned, neutral-bright, fairly smooth, average recording, quiet background, normally alert, slightly relaxed, some disfluency
(normal-paced, fairly steady, average clarity, monologue) Med mig i studio har jag är skribent och FRP-politiker Maria Seler. Hun är historisk och blir ett sistingsvald. Den första transpersonen som har valt till Storting är riktigt och som trädje vara för oss FRP.
full caption & clip details
A child masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, authoritative; average recording, quiet background; genuineness 1.7/6; vocal-burst blend 3.8/10; 15.1s, SV.
855522_00017480 · in -25.4 dBFS · gain +5.4 dB · podcast-06106
(fast, fairly steady, somewhat unclear, casual) Politiker har blivit spodd en framtidig ledstärna och är i FRP. Här är politisk nestleder i oslof
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; average recording, quiet background; genuineness 2.8/6; vocal-burst blend 5.6/10; 8.0s, SV.
855522_00018988 · in -26.2 dBFS · gain +6.2 dB · podcast-05504
(brisk, moderately variable, average clarity, casual) och bli från olikan stämtin som vara i central styra på lands möta. Det har ju
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 6.1/10; 6.0s, SV.
855522_00019780 · in -25.7 dBFS · gain +5.7 dB · podcast-05492
Embarrassment(unconstrained axis: Fear)identity +0.45 emotion 68 %   k-B1-k3 · #9

This chain comes from the one-sided rule: only Embarrassment had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Embarrassment strongly present — 0.79, higher than 79 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.20.

Nothing was asked of the other axis, and in fact Fear drifts down from 0.99 to 0.67 (-0.33), which the rule did not require.

It takes 3 clips to get there. Clip to clip the moves are +0.11, then +0.10 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.19 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.22 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.19, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 51 s · en · podcast

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.166 before conversion and 0.619 after — it rose by 0.453. Neighbour-to-neighbour the worst pair went 0.211 → 0.710. (The earlier render, with segment 1 left raw, scores 0.581 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Embarrassment moved +0.201 in the original and +0.137 after conversion — 68 % of the delta retained. On the other named axis, Fear, -0.328 became -0.720.

Quality. Mean predicted overall quality across the segments went 2.87 → 3.06 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.166 → 0.619 +0.453identity cos neighbours 0.211 → 0.710d_b rescored +0.201 → +0.137d_a rescored -0.328 → -0.720d_a mined -0.328d_b mined 0.202min_cos_consec (site) 0.2169min_cos_anchor (site) 0.1877dataset podcastlang enspeaker 1945total 50.6schain gain +5.6 dBseam step 0.8 dBcrossfades 150/100 ms
Script — 3 chunks, 3 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, fairly smooth, average recording, quiet background, neutral tension, moderately variable, some disfluency, average clarity
(fear, confusion, doubt · normal-paced, very low-energy, normal breath, casual) Angie can kick rocks. I can't with Whitney. I cannot. Her voice and like I can literally see how slow the wills turn in her head. She's like, what, Lisa? But but (surprised gasp) uh like it's just like I don't compute like I can't with her. I I can't. You know
full caption & clip details
A young adult feminine voice; delivery is very low-energy, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, slightly thin; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as fear, confusion, doubt; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 3.8/6; vocal-burst blend 3.6/10; 19.4s, EN.
1945_00829352 · in -27.2 dBFS · gain +7.2 dB · podcast-00388
(infatuation, amusement, jealousy and envy · brisk, energised, light breath, casual) realize. (ahem) Up until
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as infatuation, amusement, jealousy and envy; style: casual, conversational; average recording, quiet background; genuineness 5.5/6; vocal-burst blend 8.0/10; 18.0s, EN.
1945_00831320 · in -25.0 dBFS · gain +5.0 dB · podcast-02105
(embarrassment, doubt, pride · normal-paced, energised, light breath, casual) I'm like, oh, more, please. But (low mumble) um I I think I figured it out, at least in my mind. Like I know, Candy, you have always said how the other girls are very jealous of her, which I believe is true also.
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as embarrassment, doubt, pride; style: casual, conversational; average recording, quiet background; genuineness 5.0/6; vocal-burst blend 6.5/10; 13.4s, EN.
1945_00833424 · in -25.4 dBFS · gain +5.5 dB · podcast-05846
Anger(unconstrained axis: Pride)identity +0.15 emotion 42 %   k-B1-k3 · #10

This chain comes from the one-sided rule: only Anger had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Anger strongly present — 0.76, higher than 76 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.24.

Nothing was asked of the other axis, and in fact Pride drifts down from 0.99 to 0.91 (-0.08), which the rule did not require.

It takes 3 clips to get there. Clip to clip the moves are +0.03, then +0.20 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.19 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.19 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 34 s · fr · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.330 before conversion and 0.480 after — it rose by 0.151. Neighbour-to-neighbour the worst pair went 0.330 → 0.480. (The earlier render, with segment 1 left raw, scores 0.436 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Anger moved +0.236 in the original and +0.098 after conversion — 42 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Pride, -0.079 became -0.357.

Quality. Mean predicted overall quality across the segments went 2.72 → 3.06 (+0.33) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.330 → 0.480 +0.151identity cos neighbours 0.330 → 0.480d_b rescored +0.236 → +0.098d_a rescored -0.079 → -0.357d_a mined -0.079d_b mined 0.236min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang frspeaker FR_845xJJDmQ9gtotal 33.5schain gain +1.7 dBseam step 2.3 dBcrossfades 150/150 ms
Script — 3 chunks, 3 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, quiet background, normally alert, some disfluency
(pride, triumph, sourness · brisk, neutral tension, moderately variable, conversational) (low mumble) euh, euh, euh, euh, euh,
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, slightly guarded; reads as pride, triumph, sourness; style: conversational, casual; average recording, quiet background; genuineness 4.9/6; vocal-burst blend 10.0/10; 16.9s, FR.
FR_845xJJDmQ9g_W000020 · in -14.4 dBFS · gain -5.6 dB · emolia-02871
(confusion · normal-paced, slightly relaxed, fairly steady, casual) (low mumble) l'un de moi impacte bien, sur le plan, (ahem) (ahem) euh, sur le plan économique.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, dark, slightly rough, thin; somewhat unclear, some disfluency, moderate pitch range, audible breath; affect is neutral, slightly dominant, slightly guarded; reads as confusion; style: casual, conversational; below-average recording, quiet background; genuineness 4.9/6; vocal-burst blend 3.8/10; 4.4s, FR.
FR_845xJJDmQ9g_W000021 · in -15.4 dBFS · gain -4.6 dB · emolia-02871
(anger, confusion, malevolence malice · normal-paced, neutral tension, moderately variable, conversational) sur la table. Nous, on a des amis qui étaient là-bas, comme Birahim Sek. Y'a pas de loi qui a encore été discutée. Y'a pas de clé de répartition qui a encore été (low mumble) définie.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as anger, confusion, malevolence malice; style: conversational, authoritative; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 4.2/10; 12.7s, FR.
FR_845xJJDmQ9g_W000023 · in -15.8 dBFS · gain -4.2 dB · emolia-02871
Infatuation(unconstrained axis: Helplessness)identity −0.07 emotion 41 %   k-B1-k3 · #11

This chain comes from the one-sided rule: only Infatuation had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Infatuation clearly present — 0.60, higher than 60 % of clips in this corpus — and ends with it at the very top of the corpus at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.33.

Nothing was asked of the other axis, and in fact Helplessness drifts down from 0.96 to 0.76 (-0.20), which the rule did not require.

It takes 3 clips to get there. Clip to clip the moves are +0.24, then +0.10 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.87 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.87 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.87. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 30 s · en · podcast

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.808 before conversion and 0.739 after — it fell by 0.069. Neighbour-to-neighbour the worst pair went 0.808 → 0.739. (The earlier render, with segment 1 left raw, scores 0.760 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Infatuation moved +0.362 in the original and +0.147 after conversion — 41 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Helplessness, -0.202 became -0.031.

Quality. Mean predicted overall quality across the segments went 2.91 → 3.13 (+0.21) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.808 → 0.739 -0.069identity cos neighbours 0.808 → 0.739d_b rescored +0.362 → +0.147d_a rescored -0.202 → -0.031d_a mined -0.202d_b mined 0.332min_cos_consec (site) 0.8734min_cos_anchor (site) 0.8734dataset podcastlang enspeaker 931638total 29.3schain gain +0.5 dBseam step 2.8 dBcrossfades 100/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, quiet background, normal-paced, normally alert
(thankfulness gratitude, helplessness, emotional numbness · slightly relaxed, fairly steady, average clarity, casual) it doesn't really matter what courses you've done, like I'm doing NLP now, and you'll always be who you are as a coach.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as thankfulness gratitude, helplessness, emotional numbness; style: casual, conversational; average recording, quiet background; genuineness 4.7/6; vocal-burst blend 3.3/10; 5.9s, EN.
931638_00174912 · in -25.6 dBFS · gain +5.6 dB · podcast-02398
(sourness, anger, contempt · neutral tension, moderately variable, somewhat unclear, casual) That's why I don't mind people say the market's saturated, I'm like, not even close. I don't mind as well. I've had (low mumble) um fellow life coaches on my podcast. I'm gonna get more on. I don't mind standing shoulder to shoulder to them. I mean, you're a coach as well in a different aspect, but yeah,
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as sourness, anger, contempt; style: casual, conversational; average recording, quiet background; genuineness 5.5/6; vocal-burst blend 9.0/10; 16.9s, EN.
931638_00175512 · in -26.2 dBFS · gain +6.2 dB · podcast-04522
(infatuation, amusement · slightly relaxed, fairly steady, somewhat unclear, casual) competition because whoever wants to be coached by me will be coached by me because they'll relate to me, you know. I
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is neutral, neutral stance, neutral openness; reads as infatuation, amusement; style: casual, conversational; average recording, quiet background; genuineness 5.0/6; vocal-burst blend 2.7/10; 6.8s, EN.
931638_00177424 · in -28.6 dBFS · gain +8.6 dB · podcast-02411
Sourness(unconstrained axis: Bitterness)identity +0.13 emotion 75 %   k-B1-k3 · #12

This chain comes from the one-sided rule: only Sourness had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Sourness strongly present — 0.76, higher than 76 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.24.

Nothing was asked of the other axis, and in fact Bitterness climbs from 0.94 to 0.99 (+0.05), which the rule did not require.

It takes 3 clips to get there. Clip to clip the moves are +0.21, then +0.03 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.78 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.83 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.78, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 60 s · en · podcast

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.677 before conversion and 0.807 after — it rose by 0.130. Neighbour-to-neighbour the worst pair went 0.675 → 0.794. (The earlier render, with segment 1 left raw, scores 0.513 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Sourness moved +0.245 in the original and +0.184 after conversion — 75 % of the delta retained, which is most of it. On the other named axis, Bitterness, +0.058 became +0.183.

Quality. Mean predicted overall quality across the segments went 2.85 → 3.29 (+0.44) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.677 → 0.807 +0.130identity cos neighbours 0.675 → 0.794d_b rescored +0.245 → +0.184d_a rescored +0.058 → +0.183d_a mined 0.054d_b mined 0.237min_cos_consec (site) 0.8253min_cos_anchor (site) 0.7830dataset podcastlang enspeaker 298210total 59.2schain gain +4.0 dBseam step 1.1 dBcrossfades 150/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · slightly cool, thin, moderately variable
(bitterness, anger, pride · brisk, energised, neutral tension, cartoonish) just so you understand how it is we're going to cover when we start studying on the 21st okay what God is saying first of all that prophet Hosea was a prophet that was called during an inconvenient time much like the prophets of the old testament so much so he was going through so much that he was called while his wife was committing adultery so not only did he have to deliver God's word
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, thin; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, fairly guarded; reads as bitterness, anger, pride; style: cartoonish, dramatic; below-average recording, some background noise; genuineness 2.9/6; vocal-burst blend 6.5/10; 22.9s, EN.
298210_00451512 · in -28.0 dBFS · gain +8.0 dB · podcast-02270
(bitterness, anger, disappointment · brisk, energised, neutral tension, dramatic) okay he had to deliver it while he had other fleshly things on his mind that were just you know that that that were messing with him you know here you are trying to be a prophet to the to to God and and (surprised gasp) a prophet from God and at the same time your wife is out committing adultery and that scripture when he said he said my people will perish through the lack of knowledge what God is saying is not directed at the saints God was saying that to the priest
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, rough, thin; average clarity, some disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, fairly guarded; reads as bitterness, anger, disappointment; style: dramatic, monologue; below-average recording, quiet background; genuineness 3.0/6; vocal-burst blend 7.7/10; 26.9s, EN.
298210_00453802 · in -26.6 dBFS · gain +6.6 dB · podcast-02262
(impatience and irritability, sourness, pride · fast, highly aroused, slightly relaxed, cartoonish) because he was telling them because you failed to go ahead and train equip and disciple my people
full caption & clip details
A child masculine voice; delivery is highly aroused, fast, slightly relaxed, moderately variable; timbre is slightly cool, very dark, rough, thin; slurred, frequent disfluency, very wide pitch range, audible breath; affect is elated, neutral stance, guarded; reads as impatience and irritability, sourness, pride; style: cartoonish, dramatic; average recording, quiet background; genuineness 1.2/6; vocal-burst blend 0.8/10; 9.8s, EN.
298210_00456486 · in -30.9 dBFS · gain +10.8 dB · podcast-02283
Infatuation(unconstrained axis: Contemplation)identity −0.02 emotion 173 %   k-B1-k3 · #13

This chain comes from the one-sided rule: only Infatuation had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Infatuation clearly present — 0.70, higher than 70 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.27.

Nothing was asked of the other axis, and in fact Contemplation barely moves at all, sitting near 0.98 throughout.

It takes 3 clips to get there. Clip to clip the moves are +0.17, then +0.10 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.87 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.87 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 33 s · en · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.750 before conversion and 0.729 after — it fell by 0.022. Neighbour-to-neighbour the worst pair went 0.810 → 0.776. (The earlier render, with segment 1 left raw, scores 0.694 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Infatuation moved +0.274 in the original and +0.474 after conversion — 173 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Contemplation, -0.021 became -0.024.

Quality. Mean predicted overall quality across the segments went 2.95 → 3.18 (+0.23) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.750 → 0.729 -0.022identity cos neighbours 0.810 → 0.776d_b rescored +0.274 → +0.474d_a rescored -0.021 → -0.024d_a mined -0.021d_b mined 0.274min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_B00059_S01107total 32.1schain gain -0.6 dBseam step 4.2 dBcrossfades 150/100 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a middle-aged masculine voice · neutral-toned, balanced body, slightly relaxed, fairly steady
(contemplation, distress, disappointment · slow, very low-energy, frequent disfluency, monologue) Real problems with not just how we talk about Jesus, but even believing
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, slow, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly negative, neutral stance, neutral openness; reads as contemplation, distress, disappointment; style: monologue, whispered; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 1.8/10; 5.9s, EN.
EN_B00059_S01107_W000015 · in -13.9 dBFS · gain -6.1 dB · emolia-01383
(awe, contemplation · measured, normally alert, some disfluency, casual) Right? So that Jesus, Jesus is going around doing miracles, kinda with the same power the apostles would have later.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as awe, contemplation; style: casual, monologue; good recording, quiet background; genuineness 2.1/6; vocal-burst blend 1.7/10; 8.4s, EN.
EN_B00059_S01107_W000016 · in -13.1 dBFS · gain -6.9 dB · emolia-01383
(infatuation, concentration, contemplation · measured, normally alert, some disfluency, whispered) Or, no, he's not really a person. No, we want to, we want to leave both intact when we say the blood of Jesus purifies us from all sins. We confess his full human nature and his full divine nature and say that they are both at work contributing the unique attributes to the person, right?
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, normal breath; affect is mildly negative, slightly dominant, slightly guarded; reads as infatuation, concentration, contemplation; style: whispered, monologue; average recording, quiet background; genuineness 1.5/6; vocal-burst blend 2.9/10; 18.3s, EN.
EN_B00059_S01107_W000017 · in -13.9 dBFS · gain -6.1 dB · emolia-01383
Pride(unconstrained axis: Anger)identity +0.54 emotion 232 %   k-B1-k3 · #14

This chain comes from the one-sided rule: only Pride had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Pride clearly present — 0.68, higher than 68 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.31.

Nothing was asked of the other axis, and in fact Anger barely moves at all, sitting near 0.95 throughout.

It takes 3 clips to get there. Clip to clip the moves are +0.17, then +0.14 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.25 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.25 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.25, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 55 s · en · podcast

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.302 before conversion and 0.840 after — it rose by 0.539. Neighbour-to-neighbour the worst pair went 0.302 → 0.840. (The earlier render, with segment 1 left raw, scores 0.736 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Pride moved +0.311 in the original and +0.722 after conversion — 232 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Anger, +0.000 became +0.169.

Quality. Mean predicted overall quality across the segments went 2.84 → 3.28 (+0.45) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.302 → 0.840 +0.539identity cos neighbours 0.302 → 0.840d_b rescored +0.311 → +0.722d_a rescored +0.000 → +0.169d_a mined -0.000d_b mined 0.310min_cos_consec (site) 0.2500min_cos_anchor (site) 0.2500dataset podcastlang enspeaker 829953total 54.3schain gain +1.5 dBseam step 1.1 dBcrossfades 100/100 ms
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, neutral tension, some disfluency, light breath
(anger, disgust, interest · normal-paced, normally alert, moderately variable, casual) all the inversion that has converting in a ROI null, if it's negative.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as anger, disgust, interest; style: casual, conversational; average recording, quiet background; genuineness 5.0/6; vocal-burst blend 10.0/10; 11.4s, EN.
829953_00028244 · in -23.3 dBFS · gain +3.3 dB · podcast-03465
(hope enthusiasm optimism, interest, pleasure ecstasy · brisk, energised, moderately variable, casual) Correcto. Va un poco encadenada ahí. Te puede dar innovación, te puede dar actualidad, ¿no? Hoy en día, (ahem) de pronto tener una application se vuelve como esto que está este como en boga, ¿no? Como en su momento fue el tener una página web. Pero no específicamente tal vez es una application, ¿no? Que es justamente el tema que tocabas. ¿Por qué yo la quería? Pensemos, y es lo que te lo voy a regresar con una
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as hope enthusiasm optimism, interest, pleasure ecstasy; style: casual, dramatic; below-average recording, quiet background; genuineness 5.0/6; vocal-burst blend 10.0/10; 22.2s, EN.
829953_00037360 · in -24.8 dBFS · gain +4.8 dB · podcast-01978
(pride, anger, jealousy and envy · brisk, normally alert, fairly steady, monologue) pregunta. ¿Por qué yo quiero tener una application? ¿Qué me da una application que no me dé un web app? No, por ejemplo, me da personalización. ¿Por qué? Porque lo que yo tengo in my dispositivo sale es único para mí, ¿no? Yo tengo un celular de cierto color, con ciertas características, y con ciertas (low mumble)
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as pride, anger, jealousy and envy; style: monologue, authoritative; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 8.6/10; 21.0s, EN.
829953_00039584 · in -20.2 dBFS · gain +0.2 dB · podcast-03459
Concentration(unconstrained axis: Hope Enthusiasm Optimism)identity +0.01 emotion 103 %   k-B1-k3 · #15

This chain comes from the one-sided rule: only Concentration had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Concentration clearly present — 0.73, higher than 73 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.26.

Nothing was asked of the other axis, and in fact Hope Enthusiasm Optimism drifts down from 0.97 to 0.82 (-0.15), which the rule did not require.

It takes 3 clips to get there. Clip to clip the moves are +0.14, then +0.12 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? Not directly measured. What does exist is a timbre similarity of 0.86 against the first clip, which describes how alike the voices sound rather than whether they are the same person. It sits on a different scale from the identity check (corpus-wide the timbre numbers run much higher), so it cannot be read against the 0.80 identity threshold and is given here without a pass or fail.

Voice consistency: these clips are separate recordings joined together. Speaker identity was not measured for this sample; the available timbre similarity of 0.86 says the voices sound broadly alike but is not a same-person check. You may notice the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 53 s · en · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.860 before conversion and 0.868 after — it rose by 0.008. Neighbour-to-neighbour the worst pair went 0.860 → 0.868. (The earlier render, with segment 1 left raw, scores 0.654 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.261 in the original and +0.268 after conversion — 103 % of the delta retained, which is essentially all of it. On the other named axis, Hope Enthusiasm Optimism, -0.151 became -0.231.

Quality. Mean predicted overall quality across the segments went 2.39 → 3.20 (+0.81) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.860 → 0.868 +0.008identity cos neighbours 0.860 → 0.868d_b rescored +0.261 → +0.268d_a rescored -0.151 → -0.231d_a mined -0.152d_b mined 0.261min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_r3Qi2hymMF8total 52.5schain gain +3.3 dBseam step 1.4 dBcrossfades 100/100 ms
Script — 3 chunks, 3 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, fairly smooth, balanced body, average recording, quiet background, normally alert, slightly relaxed, fairly steady
(hope enthusiasm optimism, elation, interest · brisk, casual, playful) Yeah, absolutely. So with the launch of our new product, (ahem) uh, Docker Enterprise Container Cloud, we're making two subscriptions available as well, uh, (low mumble) (ahem) um, named Prod Care, which is a 24 seven mission critical support offering.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as hope enthusiasm optimism, elation, interest; style: casual, playful; average recording, quiet background; genuineness 2.6/6; vocal-burst blend 2.2/10; 13.2s, EN.
EN_r3Qi2hymMF8_W000007 · in -16.5 dBFS · gain -3.5 dB · emolia-02194
(interest, pride, hope enthusiasm optimism · normal-paced, monologue) And, (ahem) uh, Opscare being a fully managed platform as a service, uh, (low mumble) um, subscription. Now these, these offerings have been available on the Mirantis cloud platform side of the, of our, of our business for quite some time. We've been very successful with them. So it's really excited making them available to our Docker enterprise customers. (low mumble) Um, so.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as interest, pride, hope enthusiasm optimism; style: monologue; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 3.0/10; 20.8s, EN.
EN_r3Qi2hymMF8_W000008 · in -17.5 dBFS · gain -2.5 dB · emolia-02194
(concentration, interest · normal-paced, casual, monologue) What we, what we're trying to achieve with these accounts are (low mumble) with the, with these, uh, (ahem) subscriptions rather, you know, (ahem) uh, 30% of the, the Fortune 100 companies are Mirantis, (ahem) uhm, uh, customers. So we will work on a day-to-day basis with them with their container and Kubernetes initiatives. So when we speak to these customers,
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as concentration, interest; style: casual, monologue; average recording, quiet background; genuineness 2.7/6; vocal-burst blend 2.0/10; 18.8s, EN.
EN_r3Qi2hymMF8_W000009 · in -17.4 dBFS · gain -2.6 dB · emolia-02194
Fear(unconstrained axis: Pride)identity −0.16 emotion 53 %   k-B1-k3 · #16

This chain comes from the one-sided rule: only Fear had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Fear clearly present — 0.73, higher than 73 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.26.

Nothing was asked of the other axis, and in fact Pride drifts down from 0.77 to 0.24 (-0.53), which the rule did not require.

It takes 3 clips to get there. Clip to clip the moves are +0.23, then +0.03 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 11 s · en · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.783 before conversion and 0.626 after — it fell by 0.157. Neighbour-to-neighbour the worst pair went 0.819 → 0.679. (The earlier render, with segment 1 left raw, scores 0.620 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fear moved +0.264 in the original and +0.139 after conversion — 53 % of the delta retained. On the other named axis, Pride, -0.534 became +0.222.

Quality. Mean predicted overall quality across the segments went 2.37 → 2.58 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.783 → 0.626 -0.157identity cos neighbours 0.819 → 0.679d_b rescored +0.264 → +0.139d_a rescored -0.534 → +0.222d_a mined -0.534d_b mined 0.264min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang enspeaker EN_wTo0-i_xavQtotal 10.3schain gain +1.8 dBseam step 1.3 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, balanced body, average recording, quiet background, normally alert, slightly relaxed, fairly steady, some disfluency
(normal-paced, light breath, casual, conversational) And could potentially turn one into a zombie
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, conversational; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 2.3/10; 3.2s, EN.
EN_wTo0-i_xavQ_W000004 · in -16.3 dBFS · gain -3.7 dB · emolia-00717
(fear, helplessness, distress · normal-paced, light breath, casual, authoritative) The fungus essentially feeds on the brain of the victim.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as fear, helplessness, distress; style: casual, authoritative; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 1.2/10; 3.5s, EN.
EN_wTo0-i_xavQ_W000006 · in -15.4 dBFS · gain -4.6 dB · emolia-00717
(fear, disgust, distress · measured, normal breath, casual, storytelling) And since more disturbing than the scariest horror film.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, normal breath; affect is neutral, slightly dominant, neutral openness; reads as fear, disgust, distress; style: casual, storytelling; average recording, quiet background; genuineness 2.4/6; vocal-burst blend 2.0/10; 4.0s, EN.
EN_wTo0-i_xavQ_W000008 · in -15.9 dBFS · gain -4.1 dB · emolia-00717
Hope Enthusiasm Optimism(unconstrained axis: Contemplation)identity −0.02 emotion 70 %   k-B1-k3 · #17

This chain comes from the one-sided rule: only Hope Enthusiasm Optimism had to get where it was going, by at least 0.20. The other emotion was left completely free.

The chain starts with Hope Enthusiasm Optimism clearly present — 0.73, higher than 73 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.25.

Nothing was asked of the other axis, and in fact Contemplation drifts down from 0.92 to 0.72 (-0.19), which the rule did not require.

It takes 3 clips to get there. Clip to clip the moves are +0.13, then +0.12 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.92 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.95 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.92), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 54 s · en · podcast

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.890 before conversion and 0.866 after — it fell by 0.024. Neighbour-to-neighbour the worst pair went 0.931 → 0.861. (The earlier render, with segment 1 left raw, scores 0.648 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Hope Enthusiasm Optimism moved +0.248 in the original and +0.172 after conversion — 70 % of the delta retained. On the other named axis, Contemplation, -0.197 became -0.087.

Quality. Mean predicted overall quality across the segments went 2.86 → 3.20 (+0.35) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.890 → 0.866 -0.024identity cos neighbours 0.931 → 0.861d_b rescored +0.248 → +0.172d_a rescored -0.197 → -0.087d_a mined -0.192d_b mined 0.247min_cos_consec (site) 0.9550min_cos_anchor (site) 0.9247dataset podcastlang enspeaker 872020total 52.9schain gain +3.2 dBseam step 1.6 dBcrossfades 150/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, slightly rough, balanced body, average recording, quiet background, brisk, subdued
(contemplation · slightly tense, casual, monologue) somebody somewhere is looking at you, needs you, needs your example, needs your, you know, your positivity, your energy, in order to give them an example that they may not have had, even if you haven't had that example yourself. You know what I mean?
full caption & clip details
An adult masculine voice; delivery is subdued, brisk, slightly tense, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as contemplation; style: casual, monologue; average recording, quiet background; genuineness 4.5/6; vocal-burst blend 9.1/10; 14.6s, EN.
872020_00504576 · in -24.0 dBFS · gain +4.0 dB · podcast-03004
(contemplation, interest, teasing · neutral tension, casual, conversational) Try to be that example if you haven't had that example. Find the example if you need an example. There's some some male somewhere around that you can glean from. Don't be don't be too proud to say, Hey man, how'd you learn how to do this? Where'd you get that? You know what I mean? You know.
full caption & clip details
An adult masculine voice; delivery is subdued, brisk, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly negative, slightly dominant, fairly guarded; reads as contemplation, interest, teasing; style: casual, conversational; average recording, quiet background; genuineness 5.1/6; vocal-burst blend 9.7/10; 15.0s, EN.
872020_00506035 · in -22.8 dBFS · gain +2.8 dB · podcast-03003
(hope enthusiasm optimism, elation, malevolence malice · neutral tension, casual, conversational) have the right attitude, have the right (ahem) uh approach and perspective, be humble, (ahem) uh you know, be teachable, be coachable, w however you want to put it, but just have the right attitude and approach towards life and people will gravitate towards you, people will respond to you. You'll see yourself get us get so much further. Nobody's an island. None of us got here because we were so great.
full caption & clip details
An adult masculine voice; delivery is subdued, brisk, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly negative, slightly dominant, fairly guarded; reads as hope enthusiasm optimism, elation, malevolence malice; style: casual, conversational; average recording, quiet background; genuineness 4.6/6; vocal-burst blend 8.4/10; 23.6s, EN.
872020_00510016 · in -22.0 dBFS · gain +2.0 dB · podcast-03957
Astonishment Surprise(unconstrained axis: Interest)identity +0.49 emotion 136 %   k-B1-k3 · #18

This chain comes from the one-sided rule: only Astonishment Surprise had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Astonishment Surprise around average — 0.52, higher than 52 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.43.

Nothing was asked of the other axis, and in fact Interest drifts down from 0.88 to 0.57 (-0.31), which the rule did not require.

It takes 3 clips to get there. Clip to clip the moves are +0.23, then +0.20 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.15 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.15 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.15, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 44 s · en · podcast

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.168 before conversion and 0.658 after — it rose by 0.490. Neighbour-to-neighbour the worst pair went 0.148 → 0.604. (The earlier render, with segment 1 left raw, scores 0.513 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Astonishment Surprise moved +0.427 in the original and +0.581 after conversion — 136 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Interest, -0.317 became -0.358.

Quality. Mean predicted overall quality across the segments went 2.62 → 3.09 (+0.48) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.168 → 0.658 +0.490identity cos neighbours 0.148 → 0.604d_b rescored +0.427 → +0.581d_a rescored -0.317 → -0.358d_a mined -0.312d_b mined 0.427min_cos_consec (site) 0.1480min_cos_anchor (site) 0.1539dataset podcastlang enspeaker 635847total 43.0schain gain +3.1 dBseam step 5.2 dBcrossfades 100/150 ms
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · balanced body, quiet background
(measured, normally alert, relaxed, casual) I wanted to bring up the movie (low mumble) uh with uh (low mumble) Mark Wahlberg, what is it called? Uh (low mumble) the infinite, right? Those that are aware of past lifetimes, right?
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, relaxed, fairly steady; timbre is slightly cool, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: casual, conversational; below-average recording, quiet background; genuineness 4.8/6; vocal-burst blend 0.0/10; 11.0s, EN.
635847_00657032 · in -21.4 dBFS · gain +1.4 dB · podcast-01185
(intoxication altered states of consciousness, fear, confusion · measured, very low-energy, relaxed, casual) (ahem) The incident. Oh, yeah. I mean, there's all kinds of stuff to explore with this. I mean, because again, you know, if you recognize there's some kind of spiritual warfare going on, and there's some higher purpose to everything we're in right now. You know, for some they go down the simulation theory path, some, you know, like me, it's flat earth. (low mumble) Um, some you know, find in various ways, but
full caption & clip details
A young adult masculine voice; delivery is very low-energy, measured, relaxed, moderately variable; timbre is neutral-toned, dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, slightly guarded; reads as intoxication altered states of consciousness, fear, confusion; style: casual, monologue; below-average recording, quiet background; genuineness 5.7/6; vocal-burst blend 5.7/10; 24.2s, EN.
635847_00658152 · in -23.4 dBFS · gain +3.4 dB · podcast-01188
(astonishment surprise, confusion · normal-paced, normally alert, neutral tension, casual) all about the Alex Jones that I heard back in 2002. That shit seems exactly what's happening right now.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as astonishment surprise, confusion; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.4/6; vocal-burst blend 3.6/10; 8.1s, EN.
635847_00660632 · in -20.2 dBFS · gain +0.2 dB · podcast-01206
Shame(unconstrained axis: Doubt)identity +0.07 emotion 166 %   k-B1-k3 · #19

This chain comes from the one-sided rule: only Shame had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Shame around average — 0.50, right about the corpus median — and ends with it at the very top of the corpus at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.43.

Nothing was asked of the other axis, and in fact Doubt drifts down from 0.91 to 0.83 (-0.08), which the rule did not require.

It takes 3 clips to get there. Clip to clip the moves are +0.22, then +0.21 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 12 s · ja · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.779 before conversion and 0.846 after — it rose by 0.067. Neighbour-to-neighbour the worst pair went 0.693 → 0.834. (The earlier render, with segment 1 left raw, scores 0.719 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Shame moved +0.428 in the original and +0.710 after conversion — 166 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Doubt, -0.080 became +0.110.

Quality. Mean predicted overall quality across the segments went 3.08 → 3.17 (+0.10) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.779 → 0.846 +0.067identity cos neighbours 0.693 → 0.834d_b rescored +0.428 → +0.710d_a rescored -0.080 → +0.110d_a mined -0.080d_b mined 0.428min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang jaspeaker JA_B00002_S01918total 11.5schain gain +0.9 dBseam step 0.7 dBcrossfades 150/100 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed
(doubt · measured, fairly steady, clear, authoritative) その仕事が終わったらもう帰ってもいいですよ。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as doubt; style: authoritative, narration; good recording, no background noise; genuineness 1.3/6; vocal-burst blend 5.1/10; 3.4s, JA.
JA_B00002_S01918_W000203 · in -21.2 dBFS · gain +1.2 dB · emolia-02979
(astonishment surprise, disgust, confusion · slow, moderately variable, crisply articulate, narration) 1 忙しそうだね。この書類僕がやってあげるよ。
full caption & clip details
An adult masculine voice; delivery is normally alert, slow, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; crisply articulate, no disfluency, moderate pitch range, audible breath; affect is mildly negative, slightly dominant, slightly guarded; reads as astonishment surprise, disgust, confusion; style: narration, storytelling; good recording, no background noise; genuineness 1.3/6; vocal-burst blend 5.0/10; 5.4s, JA.
JA_B00002_S01918_W000204 · in -23.0 dBFS · gain +3.0 dB · emolia-02979
(shame · normal-paced, fairly steady, average clarity, storytelling) もしもし、田中さんはいらっしゃいますか?
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as shame; style: storytelling, authoritative; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 4.8/10; 3.1s, JA.
JA_B00002_S01918_W000205 · in -24.1 dBFS · gain +4.0 dB · emolia-02979
Hope Enthusiasm Optimism(unconstrained axis: Astonishment Surprise)identity −0.04 emotion 37 %   k-B1-k3 · #20

This chain comes from the one-sided rule: only Hope Enthusiasm Optimism had to get where it was going, by at least 0.25. The other emotion was left completely free.

The chain starts with Hope Enthusiasm Optimism clearly present — 0.63, higher than 63 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.31.

Nothing was asked of the other axis, and in fact Astonishment Surprise drifts down from 0.96 to 0.89 (-0.07), which the rule did not require.

It takes 3 clips to get there. Clip to clip the moves are +0.23, then +0.08 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 36 s · zh · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.888 before conversion and 0.849 after — it fell by 0.039. Neighbour-to-neighbour the worst pair went 0.889 → 0.831. (The earlier render, with segment 1 left raw, scores 0.743 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Hope Enthusiasm Optimism moved +0.305 in the original and +0.112 after conversion — 37 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Astonishment Surprise, -0.065 became -0.257.

Quality. Mean predicted overall quality across the segments went 3.01 → 3.31 (+0.30) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.888 → 0.849 -0.039identity cos neighbours 0.889 → 0.831d_b rescored +0.305 → +0.112d_a rescored -0.065 → -0.257d_a mined -0.065d_b mined 0.306min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang zhspeaker ZH_B00048_S05861total 35.0schain gain +2.5 dBseam step 0.6 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a child feminine voice · neutral-bright, fairly smooth, average recording, no background noise, normally alert, slightly relaxed, moderately variable, some disfluency
(astonishment surprise · measured, dramatic, cartoonish) 火帽子果然跳的不错,大羽毛的尾巴,粉色的花环转来转去,好看极了。跳跳蛙忍不住大声叫,呱呱呱,真是棒极了。
full caption & clip details
A child feminine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as astonishment surprise; style: dramatic, cartoonish; average recording, no background noise; genuineness 1.6/6; vocal-burst blend 2.1/10; 12.2s, ZH.
ZH_B00048_S05861_W000007 · in -18.4 dBFS · gain -1.6 dB · emolia-03753
(elation · normal-paced, cartoonish, didactic) 可是火帽子跳了一会儿,也晕晕乎乎的倒下了。跳跳蛙不明白是怎么回事,红袋鼠也不明白是怎么回事。这时突然听到旁边有人说这花真香,我来尝尝味道怎么样。
full caption & clip details
A child feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; clear, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, neutral openness; reads as elation; style: cartoonish, didactic; average recording, no background noise; genuineness 2.3/6; vocal-burst blend 5.0/10; 15.5s, ZH.
ZH_B00048_S05861_W000008 · in -19.8 dBFS · gain -0.2 dB · emolia-03753
(hope enthusiasm optimism · fast, cartoonish, storytelling) 我想起来了,老师说过,花瓣和树叶不能乱吃,乱尝闻的时间太长,也会头晕的。
full caption & clip details
A child feminine voice; delivery is normally alert, fast, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; clear, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as hope enthusiasm optimism; style: cartoonish, storytelling; average recording, no background noise; genuineness 2.7/6; vocal-burst blend 6.2/10; 7.7s, ZH.
ZH_B00048_S05861_W000009 · in -19.9 dBFS · gain -0.1 dB · emolia-03753