k-AB2-k3 — voice-corrected

AB2 at chain length k=3, all corpora, at the mining floor.

This is the voice-conversion-corrected version of this tier. The original, uncorrected page is still there and unchanged: t_k-AB2-k3.html. Every card below carries both renders so you can switch between them without leaving the page. What was done, and what it measured →
20chains converted
40segments re-voiced
0.778 → 0.803median worst-to-anchor identity cosine
96 %median emotion delta retained
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken.

(underlined, plain · delivery, style) — the tag before the words. Emotions first, then how it is delivered. Underlined descriptors are the ones that change across this chain — anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words are a real non-speech sound, printed where it happens.

The full generated caption for any clip is under “full caption & clip details”. Its perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.

The tags and captions describe the ORIGINAL clips — they are the corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the converted audio are the ones printed in each card's own paragraph and chips.
Fatigue Exhaustion ↓  /  Infatuationidentity −0.03 emotion 129 %   k-AB2-k3 · #1

This chain comes from the two-sided rule: it only counts if both emotions move — Fatigue Exhaustion down and Infatuation up — by at least 0.25 each.

The chain starts with Infatuation around average — 0.54, higher than 54 % of clips in this corpus — and ends with it strongly present at 0.85, higher than 85 % of clips in this corpus. That is a total rise of 0.31.

At the same time Fatigue Exhaustion goes the other way, from 0.87 (higher than 87 % of clips in this corpus) to 0.57 (higher than 57 % of clips in this corpus), a change of -0.30. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.09, then +0.23 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.89 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.88 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.89. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 20 s · zh · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.891 before conversion and 0.864 after — it fell by 0.027. Neighbour-to-neighbour the worst pair went 0.871 → 0.855. (The earlier render, with segment 1 left raw, scores 0.816 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Infatuation moved +0.314 in the original and +0.405 after conversion — 129 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Fatigue Exhaustion, -0.296 became -0.318.

Quality. Mean predicted overall quality across the segments went 3.05 → 3.18 (+0.13) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.891 → 0.864 -0.027identity cos neighbours 0.871 → 0.855d_b rescored +0.314 → +0.405d_a rescored -0.296 → -0.318d_a mined -0.303d_b mined 0.314min_cos_consec (site) 0.8772min_cos_anchor (site) 0.8898dataset emolialang zhspeaker ZH_B00045_S02490total 19.0schain gain +3.4 dBseam step 1.3 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, balanced body, no background noise, measured, normally alert, slightly relaxed, no disfluency, clear
(fairly steady, moderate pitch range, light breath, formal) 产品开发成功,并不表示能在市场上取代喷雾干燥法生产的咖啡。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, narration; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 0.6/10; 6.8s, ZH.
ZH_B00045_S02490_W000043 · in -21.1 dBFS · gain +1.1 dB · emolia-03722
(fairly steady, moderate pitch range, light breath, monologue) 为了避免失败,通用公司开始进行商品试销,以了解各地市场的反应。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, narration; average recording, no background noise; genuineness 1.3/6; vocal-burst blend 3.3/10; 7.0s, ZH.
ZH_B00045_S02490_W000044 · in -21.3 dBFS · gain +1.3 dB · emolia-03722
(steady, fairly narrow pitch, audible breath, didactic) 这已是消息时间长达四年之久。换句话说。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; clear, no disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, formal; average recording, no background noise; genuineness 1.8/6; vocal-burst blend 0.3/10; 5.6s, ZH.
ZH_B00045_S02490_W000045 · in -22.5 dBFS · gain +2.5 dB · emolia-03722
Malevolence Malice ↓  /  Intoxication Altered States of Consciousnessidentity +0.10 emotion 137 %   k-AB2-k3 · #2

This chain comes from the two-sided rule: it only counts if both emotions move — Malevolence Malice down and Intoxication Altered States of Consciousness up — by at least 0.25 each.

The chain starts with Intoxication Altered States of Consciousness clearly present — 0.58, higher than 58 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.41.

At the same time Malevolence Malice goes the other way, from 0.72 (higher than 72 % of clips in this corpus) to 0.47 (lower than 53 % of clips in this corpus), a change of -0.25. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.23, then +0.17 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.74 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.75 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.74, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 16 s · zh · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.695 before conversion and 0.795 after — it rose by 0.100. Neighbour-to-neighbour the worst pair went 0.751 → 0.798. (The earlier render, with segment 1 left raw, scores 0.647 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Intoxication Altered States of Consciousness moved +0.407 in the original and +0.556 after conversion — 137 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Malevolence Malice, -0.253 became -0.172.

Quality. Mean predicted overall quality across the segments went 2.90 → 3.02 (+0.12) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.695 → 0.795 +0.100identity cos neighbours 0.751 → 0.798d_b rescored +0.407 → +0.556d_a rescored -0.253 → -0.172d_a mined -0.253d_b mined 0.407min_cos_consec (site) 0.7485min_cos_anchor (site) 0.7366dataset emolialang zhspeaker ZH_B00042_S04983total 15.0schain gain +1.9 dBseam step 1.8 dBcrossfades 100/100 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, balanced body, no background noise, slightly relaxed
(normal-paced, normally alert, fairly steady, formal) 我不知道这个隧道里面有没有信号,应该是有吧。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 2.5/6; vocal-burst blend 3.2/10; 3.1s, ZH.
ZH_B00042_S04983_W000024 · in -23.1 dBFS · gain +3.1 dB · emolia-03695
(emotional numbness · measured, subdued, fairly steady, ASMR) 到现在都不敢相信这么一个老顽同事的。
full caption & clip details
A young adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, neutral openness; reads as emotional numbness; style: ASMR, monologue; good recording, no background noise; genuineness 3.5/6; vocal-burst blend 3.0/10; 5.3s, ZH.
ZH_B00042_S04983_W000025 · in -24.7 dBFS · gain +4.7 dB · emolia-03695
(intoxication altered states of consciousness, confusion · slow, subdued, steady, whispered) 一位可敬可爱可亲的,因为尊长真的走。
full caption & clip details
A young adult masculine voice; delivery is subdued, slow, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as intoxication altered states of consciousness, confusion; style: whispered, ASMR; average recording, no background noise; genuineness 2.4/6; vocal-burst blend 3.4/10; 7.0s, ZH.
ZH_B00042_S04983_W000026 · in -23.3 dBFS · gain +3.3 dB · emolia-03695
Concentration ↓  /  Painidentity −0.07 emotion 96 %   k-AB2-k3 · #3

This chain comes from the two-sided rule: it only counts if both emotions move — Concentration down and Pain up — by at least 0.25 each.

The chain starts with Pain around average — 0.54, higher than 54 % of clips in this corpus — and ends with it strongly present at 0.86, higher than 86 % of clips in this corpus. That is a total rise of 0.32.

At the same time Concentration goes the other way, from 0.90 (higher than 90 % of clips in this corpus) to 0.42 (lower than 58 % of clips in this corpus), a change of -0.48. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.17, then +0.14 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.78 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.73 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.78, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 20 s · en · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.690 before conversion and 0.616 after — it fell by 0.074. Neighbour-to-neighbour the worst pair went 0.693 → 0.715. (The earlier render, with segment 1 left raw, scores 0.519 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Pain moved +0.317 in the original and +0.305 after conversion — 96 % of the delta retained, which is essentially all of it. On the other named axis, Concentration, -0.477 became -0.583.

Quality. Mean predicted overall quality across the segments went 2.29 → 2.81 (+0.52) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.690 → 0.616 -0.074identity cos neighbours 0.693 → 0.715d_b rescored +0.317 → +0.305d_a rescored -0.477 → -0.583d_a mined -0.478d_b mined 0.317min_cos_consec (site) 0.7273min_cos_anchor (site) 0.7831dataset emolialang enspeaker EN_y9YTYNnIrqQtotal 19.0schain gain +3.9 dBseam step 1.4 dBcrossfades 100/150 ms
Script — 3 chunks, 3 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, quiet background, normally alert, average clarity
(normal-paced, neutral tension, moderately variable, casual) Caste are there, ethnicity is divided into caste in India, but linguistic divisions are there when, uh, (low mumble) states were constituted in India in 1950 and after.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; no dominant emotion; style: casual, monologue; average recording, quiet background; genuineness 4.5/6; vocal-burst blend 6.6/10; 10.0s, EN.
EN_y9YTYNnIrqQ_W000169 · in -17.7 dBFS · gain -2.3 dB · emolia-02376
(brisk, slightly relaxed, fairly steady, casual) Of course, different, (ahem) uh, slangs are there in different parts of Kerala, but mostly the same language.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, playful; average recording, quiet background; genuineness 4.4/6; vocal-burst blend 2.8/10; 6.0s, EN.
EN_y9YTYNnIrqQ_W000170 · in -17.3 dBFS · gain -2.7 dB · emolia-02376
(measured, slightly relaxed, fairly steady, casual) A (ahem) sect of people who joined, (low mumble) uh, to the
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, formal; average recording, quiet background; genuineness 2.8/6; vocal-burst blend 0.8/10; 3.3s, EN.
EN_y9YTYNnIrqQ_W000171 · in -17.2 dBFS · gain -2.8 dB · emolia-02376
Impatience and Irritability ↓  /  Hope Enthusiasm Optimismidentity +0.32 emotion 9 %   k-AB2-k3 · #4

This chain comes from the two-sided rule: it only counts if both emotions move — Impatience and Irritability down and Hope Enthusiasm Optimism up — by at least 0.25 each.

The chain starts with Hope Enthusiasm Optimism clearly present — 0.66, higher than 66 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.33.

At the same time Impatience and Irritability goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.61 (higher than 61 % of clips in this corpus), a change of -0.36. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.11, then +0.22 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.49 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.46 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.49, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 47 s · en · podcast

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.495 before conversion and 0.816 after — it rose by 0.321. Neighbour-to-neighbour the worst pair went 0.439 → 0.741. (The earlier render, with segment 1 left raw, scores 0.654 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Hope Enthusiasm Optimism moved +0.319 in the original and +0.030 after conversion — 9 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Impatience and Irritability, -0.358 became -0.209.

Quality. Mean predicted overall quality across the segments went 2.69 → 3.17 (+0.48) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.495 → 0.816 +0.321identity cos neighbours 0.439 → 0.741d_b rescored +0.319 → +0.030d_a rescored -0.358 → -0.209d_a mined -0.360d_b mined 0.326min_cos_consec (site) 0.4647min_cos_anchor (site) 0.4908dataset podcastlang enspeaker 68081total 46.3schain gain +2.6 dBseam step 0.8 dBcrossfades 100/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, average recording, quiet background, neutral tension, moderately variable, wide pitch range
(impatience and irritability, confusion, shame · normal-paced, energised, some disfluency, casual) big like staples. And I go, get out is the one that I go, we'll always remember get out, we'll always talk about get out. And I go, like, us totally has that thing. We are like that. I don't know that I feel that's such an individual. I'm
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as impatience and irritability, confusion, shame; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.7/6; vocal-burst blend 9.1/10; 11.4s, EN.
68081_00066864 · in -24.5 dBFS · gain +4.5 dB · podcast-01670
(infatuation, embarrassment, doubt · brisk, normally alert, some disfluency, casual) happy you made something weird. Like, yeah, I am too, but I just don't know that I have that exact sentiment. Like I I didn't think of it as a film from this year. Like I didn't really think of it as a film.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as infatuation, embarrassment, doubt; style: casual, conversational; average recording, quiet background; genuineness 5.1/6; vocal-burst blend 8.6/10; 9.2s, EN.
68081_00068008 · in -25.7 dBFS · gain +5.7 dB · podcast-01655
(hope enthusiasm optimism, contentment, affection · normal-paced, very low-energy, frequent disfluency, casual) A (wistful sigh) lot, be as liberal as he wants.
full caption & clip details
A young adult somewhat feminine voice; delivery is very low-energy, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, slightly rough, slightly thin; somewhat unclear, frequent disfluency, wide pitch range, normal breath; affect is positive, slightly submissive, neutral openness; reads as hope enthusiasm optimism, contentment, affection; style: casual, playful; average recording, quiet background; genuineness 2.4/6; vocal-burst blend 3.7/10; 26.1s, EN.
68081_00076987 · in -25.2 dBFS · gain +5.2 dB · podcast-04972
Impatience and Irritability ↓  /  Jealousy and Envyidentity +0.05 emotion 108 %   k-AB2-k3 · #5

This chain comes from the two-sided rule: it only counts if both emotions move — Impatience and Irritability down and Jealousy and Envy up — by at least 0.25 each.

The chain starts with Jealousy and Envy around average — 0.53, higher than 53 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.41.

At the same time Impatience and Irritability goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.70 (higher than 70 % of clips in this corpus), a change of -0.27. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.19, then +0.22 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.81 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.84 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.81. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 37 s · de · podcast

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.800 before conversion and 0.853 after — it rose by 0.053. Neighbour-to-neighbour the worst pair went 0.837 → 0.869. (The earlier render, with segment 1 left raw, scores 0.572 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Jealousy and Envy moved +0.386 in the original and +0.419 after conversion — 108 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Impatience and Irritability, -0.278 became -0.334.

Quality. Mean predicted overall quality across the segments went 2.65 → 3.18 (+0.53) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.800 → 0.853 +0.053identity cos neighbours 0.837 → 0.869d_b rescored +0.386 → +0.419d_a rescored -0.278 → -0.334d_a mined -0.273d_b mined 0.408min_cos_consec (site) 0.8389min_cos_anchor (site) 0.8128dataset podcastlang despeaker 806651total 36.2schain gain +0.4 dBseam step 2.1 dBcrossfades 100/100 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, balanced body, normally alert, some disfluency, light breath
(impatience and irritability, intoxication altered states of consciousness, fatigue exhaustion · normal-paced, neutral tension, fairly steady, casual) hab zum Glück vor kurzem sogar eine Rechtsschutzversicherung abgeschlossen. Mal gucken, wie weit mir das was bringt. Aber allzu groß sehe ich meine Chancen nicht, dass ich da meine 800 Euro wieder bekommen, ehrlicherweise.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as impatience and irritability, intoxication altered states of consciousness, fatigue exhaustion; style: casual, conversational; below-average recording, some background noise; genuineness 5.7/6; vocal-burst blend 2.9/10; 13.3s, DE.
806651_00064576 · in -18.9 dBFS · gain -1.1 dB · podcast-00262
(confusion, disappointment, fatigue exhaustion · normal-paced, neutral tension, moderately variable, conversational) das hab ich mir halt auch gemacht. Also, genau, ich verstehe noch nicht so richtig, was die da gemacht haben. Die Person auf dem Perso war von 2003, also (ahem) 19, 19 Jahre alt.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, wide pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as confusion, disappointment, fatigue exhaustion; style: conversational, casual; average recording, quiet background; genuineness 5.9/6; vocal-burst blend 1.5/10; 12.2s, DE.
806651_00066256 · in -20.5 dBFS · gain +0.5 dB · podcast-04137
(jealousy and envy, confusion, sourness · brisk, slightly relaxed, fairly steady, conversational) Ja, eine Cloud Taperso, aber genau stimmt, was auch noch dazu kam. Die hat mir den Perso, also diese Bilder per Mail geschickt. Und die E-Mail-Adresse von dem Paypal-Konto war genau die gleiche, von der ich die Mail bekommen habe.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as jealousy and envy, confusion, sourness; style: conversational, playful; good recording, some background noise; genuineness 3.8/6; vocal-burst blend 3.5/10; 10.9s, DE.
806651_00068992 · in -18.9 dBFS · gain -1.1 dB · podcast-00237
Contemplation ↓  /  Reliefidentity −0.08 emotion 59 %   k-AB2-k3 · #6

This chain comes from the two-sided rule: it only counts if both emotions move — Contemplation down and Relief up — by at least 0.25 each.

The chain starts with Relief clearly present — 0.66, higher than 66 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.29.

At the same time Contemplation goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.72 (higher than 72 % of clips in this corpus), a change of -0.27. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.15, then +0.14 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.91 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.92 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.91), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 33 s · en · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.884 before conversion and 0.808 after — it fell by 0.076. Neighbour-to-neighbour the worst pair went 0.866 → 0.795. (The earlier render, with segment 1 left raw, scores 0.747 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Relief moved +0.292 in the original and +0.171 after conversion — 59 % of the delta retained. On the other named axis, Contemplation, -0.269 became -0.229.

Quality. Mean predicted overall quality across the segments went 2.78 → 3.12 (+0.34) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.884 → 0.808 -0.076identity cos neighbours 0.866 → 0.795d_b rescored +0.292 → +0.171d_a rescored -0.269 → -0.229d_a mined -0.269d_b mined 0.292min_cos_consec (site) 0.9225min_cos_anchor (site) 0.9148dataset emolialang enspeaker EN_GuZI4EqK7a4total 32.6schain gain +5.0 dBseam step 0.8 dBcrossfades 150/100 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, slightly bright, fairly smooth, balanced body, good recording, slightly relaxed, moderate pitch range
(contemplation, concentration · normal-paced, normally alert, fairly steady, casual) Is the, the body's response, the body, mind and spirits response to the traumas that you experienced earlier in your life. So in lesson and tell.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as contemplation, concentration; style: casual, monologue; good recording, quiet background; genuineness 2.3/6; vocal-burst blend 1.9/10; 8.5s, EN.
EN_GuZI4EqK7a4_W000068 · in -20.1 dBFS · gain +0.1 dB · emolia-00669
(pleasure ecstasy, contentment, concentration · normal-paced, very low-energy, steady, didactic) But to do so joyfully, to do so gratefully, to do so with optimism and hope and flow requires a different energy than one that's mostly captivated by trauma. So if you've got corporate trauma,
full caption & clip details
A young adult feminine voice; delivery is very low-energy, normal-paced, slightly relaxed, steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, minimal breath; affect is mildly positive, neutral stance, neutral openness; reads as pleasure ecstasy, contentment, concentration; style: didactic, monologue; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 0.5/10; 14.9s, EN.
EN_GuZI4EqK7a4_W000071 · in -21.0 dBFS · gain +1.0 dB · emolia-00669
(relief, concentration, fear · measured, very low-energy, steady, monologue) It's time to work with an executive coach. Again, somebody who's specifically trained and well versed in managing and resolving trauma.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, measured, slightly relaxed, steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as relief, concentration, fear; style: monologue, whispered; good recording, no background noise; genuineness 1.2/6; vocal-burst blend 0.0/10; 9.6s, EN.
EN_GuZI4EqK7a4_W000072 · in -20.8 dBFS · gain +0.8 dB · emolia-00669
Bitterness ↓  /  Interestidentity +0.01 emotion 90 %   k-AB2-k3 · #7

This chain comes from the two-sided rule: it only counts if both emotions move — Bitterness down and Interest up — by at least 0.25 each.

The chain starts with Interest clearly present — 0.70, higher than 70 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.28.

At the same time Bitterness goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.68 (higher than 68 % of clips in this corpus), a change of -0.29. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.21, then +0.07 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.92 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.91 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.92), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 46 s · en · podcast

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.888 before conversion and 0.893 after — it rose by 0.005. Neighbour-to-neighbour the worst pair went 0.867 → 0.930. (The earlier render, with segment 1 left raw, scores 0.778 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Interest moved +0.278 in the original and +0.251 after conversion — 90 % of the delta retained, which is essentially all of it. On the other named axis, Bitterness, -0.292 became -0.644.

Quality. Mean predicted overall quality across the segments went 3.07 → 3.24 (+0.17) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.888 → 0.893 +0.005identity cos neighbours 0.867 → 0.930d_b rescored +0.278 → +0.251d_a rescored -0.292 → -0.644d_a mined -0.291d_b mined 0.276min_cos_consec (site) 0.9141min_cos_anchor (site) 0.9236dataset podcastlang enspeaker 199036total 45.2schain gain +3.3 dBseam step 0.3 dBcrossfades 150/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normal-paced, normally alert, fairly steady, some disfluency
(bitterness, relief, impatience and irritability · neutral tension, casual, conversational) from a standpoint of of if you want to call it marketing, you never had to run an ad, you know, in an outdoor channel or or call me and and drop a Jimmy John commercial. You not only had me, but as soon as the the die hard hunters know, well, guess what? Jimmy is one of us because he's a hunter. (low mumble) Um,
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, slightly dominant, slightly guarded; reads as bitterness, relief, impatience and irritability; style: casual, conversational; average recording, quiet background; genuineness 4.6/6; vocal-burst blend 5.6/10; 17.8s, EN.
199036_00072472 · in -26.5 dBFS · gain +6.5 dB · podcast-03899
(relief, disappointment, awe · slightly relaxed, casual, monologue) but I had a chance to listen to Theo Vaughn's podcast and you were talking about the path and this journey. So many people not only did they not know that you hunted, but I don't think everybody understood the true story of of the success. And I I thought that you and Theo talked about it masterfully well.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as relief, disappointment, awe; style: casual, monologue; average recording, no background noise; genuineness 4.0/6; vocal-burst blend 3.8/10; 18.0s, EN.
199036_00074256 · in -28.0 dBFS · gain +8.0 dB · podcast-03890
(interest, thankfulness gratitude · neutral tension, casual, conversational) And and I wanted definitely to give some of the listeners and viewers here to s to to hear you say how how how you did begin and how how it started with the whole process of sure. Really
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as interest, thankfulness gratitude; style: casual, conversational; good recording, quiet background; genuineness 3.8/6; vocal-burst blend 7.0/10; 9.6s, EN.
199036_00076056 · in -27.9 dBFS · gain +7.9 dB · podcast-04428
Jealousy and Envy ↓  /  Doubtidentity −0.07 emotion 109 %   k-AB2-k3 · #8

This chain comes from the two-sided rule: it only counts if both emotions move — Jealousy and Envy down and Doubt up — by at least 0.25 each.

The chain starts with Doubt around average — 0.55, higher than 55 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.42.

At the same time Jealousy and Envy goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.72 (higher than 72 % of clips in this corpus), a change of -0.28. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.22, then +0.20 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.89 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.93 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.89. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 27 s · zh · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.757 before conversion and 0.685 after — it fell by 0.071. Neighbour-to-neighbour the worst pair went 0.757 → 0.711. (The earlier render, with segment 1 left raw, scores 0.585 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Doubt moved +0.419 in the original and +0.455 after conversion — 109 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Jealousy and Envy, -0.276 became -0.352.

Quality. Mean predicted overall quality across the segments went 2.87 → 3.00 (+0.13) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.757 → 0.685 -0.071identity cos neighbours 0.757 → 0.711d_b rescored +0.419 → +0.455d_a rescored -0.276 → -0.352d_a mined -0.276d_b mined 0.419min_cos_consec (site) 0.9274min_cos_anchor (site) 0.8918dataset emolialang zhspeaker ZH_B00067_S01969total 26.0schain gain +3.6 dBseam step 2.4 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · balanced body, good recording, slightly relaxed, fairly steady
(jealousy and envy, helplessness, sadness · normal-paced, normally alert, little disfluency, casual) Message jo, mustn't say too much of what we're up to.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, little disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as jealousy and envy, helplessness, sadness; style: casual, monologue; good recording, no background noise; genuineness 1.2/6; vocal-burst blend 1.4/10; 3.5s, ZH.
ZH_B00067_S01969_W000626 · in -20.4 dBFS · gain +0.3 dB · emolia-03951
(intoxication altered states of consciousness, amusement · slow, very low-energy, little disfluency, didactic) It must be done as i may say on the slide and why on the sly, i'll tell you why hip he had taken up the poker again without which.
full caption & clip details
An elderly masculine voice; delivery is very low-energy, slow, slightly relaxed, fairly steady; timbre is warm, slightly dark, slightly rough, balanced body; very clear, little disfluency, wide pitch range, normal breath; affect is mildly positive, slightly dominant, fairly guarded; reads as intoxication altered states of consciousness, amusement; style: didactic, monologue; good recording, quiet background; genuineness 0.7/6; vocal-burst blend 0.3/10; 12.6s, ZH.
ZH_B00067_S01969_W000627 · in -21.2 dBFS · gain +1.2 dB · emolia-03951
(doubt, thankfulness gratitude, malevolence malice · measured, normally alert, almost no disfluency, narration) I doubt if he could have preceded in his demonstration, your sister is given to government given to government joe.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is slightly cool, slightly dark, slightly rough, balanced body; very clear, almost no disfluency, fairly narrow pitch, normal breath; affect is neutral, slightly dominant, fairly guarded; reads as doubt, thankfulness gratitude, malevolence malice; style: narration, storytelling; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 0.0/10; 10.3s, ZH.
ZH_B00067_S01969_W000628 · in -21.6 dBFS · gain +1.6 dB · emolia-03951
Concentration ↓  /  Emotional Numbnessidentity −0.07 emotion 96 %   k-AB2-k3 · #9

This chain comes from the two-sided rule: it only counts if both emotions move — Concentration down and Emotional Numbness up — by at least 0.25 each.

The chain starts with Emotional Numbness around average — 0.50, right about the corpus median — and ends with it strongly present at 0.87, higher than 87 % of clips in this corpus. That is a total rise of 0.36.

At the same time Concentration goes the other way, from 0.84 (higher than 84 % of clips in this corpus) to 0.42 (lower than 58 % of clips in this corpus), a change of -0.42. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.21, then +0.15 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.98 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.98 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.98), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 33 s · en · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.922 before conversion and 0.855 after — it fell by 0.067. Neighbour-to-neighbour the worst pair went 0.933 → 0.888. (The earlier render, with segment 1 left raw, scores 0.608 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.361 in the original and +0.349 after conversion — 96 % of the delta retained, which is essentially all of it. On the other named axis, Concentration, -0.423 became -0.417.

Quality. Mean predicted overall quality across the segments went 2.91 → 3.05 (+0.14) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.922 → 0.855 -0.067identity cos neighbours 0.933 → 0.888d_b rescored +0.361 → +0.349d_a rescored -0.423 → -0.417d_a mined -0.423d_b mined 0.361min_cos_consec (site) 0.9779min_cos_anchor (site) 0.9779dataset emolialang enspeaker EN_KSkuPRr2Cr8total 32.3schain gain +1.6 dBseam step 1.4 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(newsreading, formal) Beginning with Sir John A. Macdonald's National Policy' 1879 and the construction of the Canadian Pacific Railway' 1875–1885 through Northern Ontario and the Canadian Prairies to British Columbia, Ontario manufacturing and industry flourished.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 15.8s, EN.
EN_KSkuPRr2Cr8_W000113 · in -15.4 dBFS · gain -4.6 dB · emolia-01830
(formal, newsreading) However, population increase slowed after a large recession hit the province in 1893, thus slowing growth drastically but for only a few years
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 8.8s, EN.
EN_KSkuPRr2Cr8_W000114 · in -14.7 dBFS · gain -5.3 dB · emolia-01830
(formal, newsreading) Many newly arrived immigrants and others moved west along the railway to the Prairie provinces and British Columbia, sparsely settling northern Ontario
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, newsreading; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.0/10; 8.2s, EN.
EN_KSkuPRr2Cr8_W000115 · in -13.8 dBFS · gain -6.2 dB · emolia-01830
Contemplation ↓  /  Fatigue Exhaustionidentity +0.04 emotion 121 %   k-AB2-k3 · #10

This chain comes from the two-sided rule: it only counts if both emotions move — Contemplation down and Fatigue Exhaustion up — by at least 0.25 each.

The chain starts with Fatigue Exhaustion clearly present — 0.59, higher than 59 % of clips in this corpus — and ends with it strongly present at 0.86, higher than 86 % of clips in this corpus. That is a total rise of 0.26.

At the same time Contemplation goes the other way, from 0.90 (higher than 90 % of clips in this corpus) to 0.57 (higher than 57 % of clips in this corpus), a change of -0.33. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.21, then +0.06 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.90 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.87 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.90), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 27 s · zh · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.861 before conversion and 0.899 after — it rose by 0.038. Neighbour-to-neighbour the worst pair went 0.879 → 0.910. (The earlier render, with segment 1 left raw, scores 0.764 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fatigue Exhaustion moved +0.288 in the original and +0.347 after conversion — 121 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Contemplation, -0.330 became -0.459.

Quality. Mean predicted overall quality across the segments went 3.12 → 3.25 (+0.13) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.861 → 0.899 +0.038identity cos neighbours 0.879 → 0.910d_b rescored +0.288 → +0.347d_a rescored -0.330 → -0.459d_a mined -0.330d_b mined 0.264min_cos_consec (site) 0.8691min_cos_anchor (site) 0.9026dataset emolialang zhspeaker ZH_B00078_S08883total 26.2schain gain +1.8 dBseam step 0.6 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, quiet background, normally alert, slightly relaxed
(normal-paced, average clarity, monologue, authoritative) 也不能说是混淆啊,就是我觉得这个理解的方向是不一样,这块儿还是要先提一嘴原著啊。但是今天各位放心,咱们肯定说的简单点,毕竟这个白骨精实在是是吧?太熟悉了这么一个事儿。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, authoritative; average recording, quiet background; genuineness 2.1/6; vocal-burst blend 7.5/10; 14.1s, ZH.
ZH_B00078_S08883_W000012 · in -19.1 dBFS · gain -0.9 dB · emolia-04055
(measured, somewhat unclear, conversational, monologue) 首先呢我要先提一个点啊,就是在我记得西游记第十九回的时候,猪八戒刚来嗯。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: conversational, monologue; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 4.7/10; 6.8s, ZH.
ZH_B00078_S08883_W000013 · in -18.2 dBFS · gain -1.8 dB · emolia-04055
(measured, somewhat unclear, monologue, authoritative) 来了之后呢,他们相当于这会儿哥儿四个嘛,唐僧孙悟空猪八戒跟白龙马。
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, authoritative; average recording, quiet background; genuineness 2.7/6; vocal-burst blend 4.5/10; 5.8s, ZH.
ZH_B00078_S08883_W000014 · in -18.1 dBFS · gain -1.9 dB · emolia-04055
Pride ↓  /  Interestidentity −0.01 emotion 26 %   k-AB2-k3 · #11

This chain comes from the two-sided rule: it only counts if both emotions move — Pride down and Interest up — by at least 0.25 each.

The chain starts with Interest clearly present — 0.61, higher than 61 % of clips in this corpus — and ends with it strongly present at 0.87, higher than 87 % of clips in this corpus. That is a total rise of 0.26.

At the same time Pride goes the other way, from 0.94 (higher than 94 % of clips in this corpus) to 0.53 (higher than 53 % of clips in this corpus), a change of -0.41. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.10, then +0.16 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.94 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.95 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.94), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 48 s · da · podcast

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.942 before conversion and 0.931 after — it fell by 0.012. Neighbour-to-neighbour the worst pair went 0.956 → 0.931. (The earlier render, with segment 1 left raw, scores 0.889 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Interest moved +0.261 in the original and +0.068 after conversion — 26 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Pride, -0.416 became -0.381.

Quality. Mean predicted overall quality across the segments went 3.12 → 3.37 (+0.25) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.942 → 0.931 -0.012identity cos neighbours 0.956 → 0.931d_b rescored +0.261 → +0.068d_a rescored -0.416 → -0.381d_a mined -0.414d_b mined 0.262min_cos_consec (site) 0.9521min_cos_anchor (site) 0.9425dataset podcastlang daspeaker 32007total 47.0schain gain +2.5 dBseam step 1.1 dBcrossfades 150/100 ms
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, good recording, quiet background, normal-paced, normally alert, slightly relaxed
(pride, helplessness, shame · fairly steady, moderate pitch range, casual, monologue) (ahem) Jeg ved ikke man hørt om ham i nyhedning, men (ahem) der er det her kunstmaser om kunst i Aalborg. (ahem) Og han (low mumble) i (ahem) feltalen skulle have det her kunstværk hængende. (ahem) Bestå af to glas rammer med en masse penge sæler i mellem. Så han har simpelthen fået en halv million kroner kontant af det her Kundsmæum.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as pride, helplessness, shame; style: casual, monologue; good recording, quiet background; genuineness 3.9/6; vocal-burst blend 7.4/10; 17.3s, DA.
32007_00317012 · in -26.8 dBFS · gain +6.8 dB · podcast-03080
(shame, sourness, impatience and irritability · moderately variable, wide pitch range, playful, conversational) Men i stedet før, så har han skabt et nyt værk med titlen Take The Money and Run, og har simpelthen bare taget penge med sig efter to tomme glas plæder. Det er simpelthen stræg simpelthen. Og jeg følger afgængt, så skal penge og læge tilbage. Men det her i et morgen og sige. Det kommer ikke til at sket.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as shame, sourness, impatience and irritability; style: playful, conversational; good recording, quiet background; genuineness 2.3/6; vocal-burst blend 1.8/10; 15.2s, DA.
32007_00318736 · in -25.7 dBFS · gain +5.7 dB · podcast-03071
(fairly steady, moderate pitch range, casual, monologue) (ahem) Han næder, og han mener ikke, at det er (ahem) en brød på nogen af. Han er københver. Kæmpe legende. Jeg ikke behøver at forklaret, hvorfor. (ahem) Så i 65, også et godt år, tænker jeg i københavn og blive før det må noget.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; good recording, quiet background; genuineness 2.8/6; vocal-burst blend 2.2/10; 14.8s, DA.
32007_00320287 · in -26.3 dBFS · gain +6.3 dB · podcast-03481
Disgust ↓  /  Astonishment Surpriseidentity −0.15 emotion 58 %   k-AB2-k3 · #12

This chain comes from the two-sided rule: it only counts if both emotions move — Disgust down and Astonishment Surprise up — by at least 0.25 each.

The chain starts with Astonishment Surprise clearly present — 0.66, higher than 66 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.32.

At the same time Disgust goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.65 (higher than 65 % of clips in this corpus), a change of -0.34. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.15, then +0.17 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.67 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.73 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.67, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 32 s · en · podcast

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.629 before conversion and 0.482 after — it fell by 0.147. Neighbour-to-neighbour the worst pair went 0.629 → 0.482. (The earlier render, with segment 1 left raw, scores 0.532 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Astonishment Surprise moved +0.455 in the original and +0.266 after conversion — 58 % of the delta retained. On the other named axis, Disgust, -0.334 became -0.218.

Quality. Mean predicted overall quality across the segments went 2.55 → 2.92 (+0.37) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.629 → 0.482 -0.147identity cos neighbours 0.629 → 0.482d_b rescored +0.455 → +0.266d_a rescored -0.334 → -0.218d_a mined -0.335d_b mined 0.321min_cos_consec (site) 0.7349min_cos_anchor (site) 0.6738dataset podcastlang enspeaker 505490total 31.5schain gain +5.1 dBseam step 1.5 dBcrossfades 100/100 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a child masculine voice · neutral-toned, fairly smooth, average recording, energised, moderately variable, average clarity, light breath
(disgust, intoxication altered states of consciousness, teasing · normal-paced, fully relaxed, no disfluency, casual) pretty much pranks people in a super elaborate ways. They
full caption & clip details
A child masculine voice; delivery is energised, normal-paced, fully relaxed, moderately variable; timbre is neutral-toned, dark, fairly smooth, thin; average clarity, no disfluency, wide pitch range, light breath; affect is mildly negative, slightly dominant, neutral openness; reads as disgust, intoxication altered states of consciousness, teasing; style: casual, playful; average recording, no background noise; explicit content; genuineness 3.2/6; vocal-burst blend 5.6/10; 3.2s, EN.
505490_00268352 · in -24.7 dBFS · gain +4.7 dB · podcast-02491
(pleasure ecstasy, amusement, hope enthusiasm optimism · brisk, slightly relaxed, some disfluency, casual) It is
full caption & clip details
A young adult masculine voice; delivery is energised, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as pleasure ecstasy, amusement, hope enthusiasm optimism; style: casual, playful; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 8.8/10; 17.4s, EN.
505490_00270432 · in -24.1 dBFS · gain +4.1 dB · podcast-00444
(astonishment surprise, amusement, impatience and irritability · brisk, neutral tension, some disfluency, casual) prank the whole internet, the whole world. (surprised gasp) And what's crazy about that is that I was on the front lines of the pranking. I'm happy to say that I that I didn't need the internet to be pranked because they got me in a
full caption & clip details
An adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, neutral openness; reads as astonishment surprise, amusement, impatience and irritability; style: casual, conversational; average recording, quiet background; genuineness 4.8/6; vocal-burst blend 8.5/10; 11.1s, EN.
505490_00272384 · in -23.7 dBFS · gain +3.7 dB · podcast-02489
Impatience and Irritability ↓  /  Doubtidentity −0.04 emotion 138 %   k-AB2-k3 · #13

This chain comes from the two-sided rule: it only counts if both emotions move — Impatience and Irritability down and Doubt up — by at least 0.25 each.

The chain starts with Doubt clearly present — 0.64, higher than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.31.

At the same time Impatience and Irritability goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.74 (higher than 74 % of clips in this corpus), a change of -0.25. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.16, then +0.15 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.88 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.87 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.88. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 27 s · en · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.841 before conversion and 0.799 after — it fell by 0.042. Neighbour-to-neighbour the worst pair went 0.841 → 0.799. (The earlier render, with segment 1 left raw, scores 0.748 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Doubt moved +0.311 in the original and +0.427 after conversion — 138 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Impatience and Irritability, -0.251 became -0.037.

Quality. Mean predicted overall quality across the segments went 2.80 → 2.96 (+0.16) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.841 → 0.799 -0.042identity cos neighbours 0.841 → 0.799d_b rescored +0.311 → +0.427d_a rescored -0.251 → -0.037d_a mined -0.251d_b mined 0.311min_cos_consec (site) 0.8670min_cos_anchor (site) 0.8817dataset emolialang enspeaker EN_B00008_S07845total 26.6schain gain +1.7 dBseam step 2.2 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, average recording, quiet background, normally alert, fairly steady, moderate pitch range
(impatience and irritability, anger, disgust · normal-paced, neutral tension, frequent disfluency, casual) He's fucked. I spar pro boxers. I know what it means with a real motherfucker that throws crazy punches. I feel the energy, I feel the speed. That's what Porino makes me do. He makes me spar with these guys that are really going to char my eyes and my defense and everything.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as impatience and irritability, anger, disgust; style: casual, conversational; average recording, quiet background; genuineness 6.0/6; vocal-burst blend 10.0/10; 16.2s, EN.
EN_B00008_S07845_W000120 · in -20.4 dBFS · gain +0.4 dB · emolia-00417
(impatience and irritability, sourness, contempt · measured, slightly relaxed, some disfluency, casual) It is risky, but fuck, you don't learn how to race Formula 1 by driving a 60.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; reads as impatience and irritability, sourness, contempt; style: casual, conversational; average recording, quiet background; genuineness 4.4/6; vocal-burst blend 2.3/10; 4.5s, EN.
EN_B00008_S07845_W000121 · in -20.1 dBFS · gain +0.1 dB · emolia-00417
(doubt · normal-paced, relaxed, frequent disfluency, casual) So you gotta go a little crazy in there and I'm sure he's gonna be better, I'm sure he's motivated, but...
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is mildly negative, neutral stance, neutral openness; reads as doubt; style: casual, conversational; average recording, quiet background; genuineness 5.1/6; vocal-burst blend 6.7/10; 6.4s, EN.
EN_B00008_S07845_W000122 · in -20.5 dBFS · gain +0.5 dB · emolia-00417
Emotional Numbness ↓  /  Concentrationidentity −0.01 emotion 70 %   k-AB2-k3 · #14

This chain comes from the two-sided rule: it only counts if both emotions move — Emotional Numbness down and Concentration up — by at least 0.25 each.

The chain starts with Concentration clearly present — 0.68, higher than 68 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.30.

At the same time Emotional Numbness goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.75 (higher than 75 % of clips in this corpus), a change of -0.25. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.13, then +0.16 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.94 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.94 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.94), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 42 s · en · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.913 before conversion and 0.903 after — it fell by 0.010. Neighbour-to-neighbour the worst pair went 0.893 → 0.872. (The earlier render, with segment 1 left raw, scores 0.817 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.297 in the original and +0.206 after conversion — 70 % of the delta retained. On the other named axis, Emotional Numbness, -0.252 became -0.296.

Quality. Mean predicted overall quality across the segments went 3.07 → 3.27 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.913 → 0.903 -0.010identity cos neighbours 0.893 → 0.872d_b rescored +0.297 → +0.206d_a rescored -0.252 → -0.296d_a mined -0.252d_b mined 0.297min_cos_consec (site) 0.9355min_cos_anchor (site) 0.9355dataset emolialang enspeaker EN_A53BUVGQdrEtotal 41.8schain gain +1.6 dBseam step 0.6 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-bright, fairly smooth, good recording, no background noise, measured, normally alert, slightly relaxed, steady
(emotional numbness · fairly narrow pitch, no audible breath, formal, newsreading) Remarks at Expo'67, Montreal, May 25, 1967, Prime Minister Pierre Elliott Trudeau famously said that being America's neighbor, "...is like sleeping with an elephant."
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is slightly cool, neutral-bright, fairly smooth, thin; clear, no disfluency, fairly narrow pitch, no audible breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 14.1s, EN.
EN_A53BUVGQdrE_W000269 · in -15.8 dBFS · gain -4.2 dB · emolia-00486
(fear, longing · moderate pitch range, light breath, formal, monologue) No matter how friendly and even-tempered the beast, if one can call it that, one is affected by every twitch and grunt
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fear, longing; style: formal, monologue; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.0/10; 8.3s, EN.
EN_A53BUVGQdrE_W000270 · in -13.6 dBFS · gain -6.4 dB · emolia-00486
(concentration · moderate pitch range, no audible breath, newsreading, formal) Prime Minister Pierre Elliott Trudeau, sharply at odds with the U.S. over Cold War policy, warned at a press conference in 1971 that the overwhelming American presence posed, a danger to our national identity from a cultural, economic and perhaps even military point of view.
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, no audible breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: newsreading, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 19.7s, EN.
EN_A53BUVGQdrE_W000271 · in -14.7 dBFS · gain -5.3 dB · emolia-00486
Teasing ↓  /  Doubtidentity +0.53 emotion 53 %   k-AB2-k3 · #15

This chain comes from the two-sided rule: it only counts if both emotions move — Teasing down and Doubt up — by at least 0.25 each.

The chain starts with Doubt clearly present — 0.64, higher than 64 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.36.

At the same time Teasing goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.74 (higher than 74 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.15, then +0.20 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.17 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.55 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.17, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 32 s · en · podcast

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.208 before conversion and 0.736 after — it rose by 0.527. Neighbour-to-neighbour the worst pair went 0.481 → 0.759. (The earlier render, with segment 1 left raw, scores 0.425 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Doubt moved +0.358 in the original and +0.190 after conversion — 53 % of the delta retained. On the other named axis, Teasing, -0.258 became -0.635.

Quality. Mean predicted overall quality across the segments went 2.80 → 3.03 (+0.23) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.208 → 0.736 +0.527identity cos neighbours 0.481 → 0.759d_b rescored +0.358 → +0.190d_a rescored -0.258 → -0.635d_a mined -0.256d_b mined 0.358min_cos_consec (site) 0.5461min_cos_anchor (site) 0.1716dataset podcastlang enspeaker 801948total 31.0schain gain +1.5 dBseam step 1.4 dBcrossfades 150/100 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, normal-paced, normally alert, some disfluency, average clarity
(teasing, affection, contempt · neutral tension, moderately variable, moderate pitch range, casual) yeah, and she's an eva's total caricature of like an uptight woman, and she's initially villainized, so like it starts off bad, like okay, this is fishy. In the middle is sort of decent because it treats Eva like a human being when she's with Ray, but
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as teasing, affection, contempt; style: casual; average recording, quiet background; mildly explicit content; genuineness 5.5/6; vocal-burst blend 9.1/10; 13.7s, EN.
801948_00517216 · in -29.1 dBFS · gain +9.1 dB · podcast-01926
(contemplation, infatuation, affection · neutral tension, moderately variable, wide pitch range, casual) and even that ends up being slightly nefarious because Eva was only a human being because she was with a man, like yeah, like they should have given her more loose personality when she when she's just with her sisters, like
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as contemplation, infatuation, affection; style: casual, conversational; average recording, quiet background; mildly explicit content; genuineness 5.3/6; vocal-burst blend 6.3/10; 13.6s, EN.
801948_00518584 · in -29.9 dBFS · gain +9.9 dB · podcast-01916
(doubt, infatuation · slightly relaxed, fairly steady, moderate pitch range, casual) it wouldn't make sense that she's more uptight around the husbands because she probably doesn't
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as doubt, infatuation; style: casual, conversational; good recording, no background noise; genuineness 4.3/6; vocal-burst blend 4.7/10; 4.0s, EN.
801948_00519952 · in -31.0 dBFS · gain +11.0 dB · podcast-02904
Hope Enthusiasm Optimism ↓  /  Interestidentity −0.05 emotion 156 %   k-AB2-k3 · #16

This chain comes from the two-sided rule: it only counts if both emotions move — Hope Enthusiasm Optimism down and Interest up — by at least 0.25 each.

The chain starts with Interest around average — 0.52, higher than 52 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.47.

At the same time Hope Enthusiasm Optimism goes the other way, from 0.68 (higher than 68 % of clips in this corpus) to 0.97 (higher than 97 % of clips in this corpus), a change of +0.29. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.23, then +0.24 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.79 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.80 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.79, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 35 s · en · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.678 before conversion and 0.623 after — it fell by 0.055. Neighbour-to-neighbour the worst pair went 0.708 → 0.639. (The earlier render, with segment 1 left raw, scores 0.536 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Interest moved +0.470 in the original and +0.732 after conversion — 156 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Hope Enthusiasm Optimism, +0.289 became +0.350.

Quality. Mean predicted overall quality across the segments went 2.36 → 2.86 (+0.50) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.678 → 0.623 -0.055identity cos neighbours 0.708 → 0.639d_b rescored +0.470 → +0.732d_a rescored +0.289 → +0.350d_a mined 0.289d_b mined 0.471min_cos_consec (site) 0.7953min_cos_anchor (site) 0.7918dataset emolialang enspeaker EN_itLhZY5lb9ytotal 34.7schain gain +5.3 dBseam step 0.2 dBcrossfades 150/100 ms
Script — 3 chunks, 3 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, fairly smooth, balanced body, normal-paced, normally alert, slightly relaxed, fairly steady, some disfluency
(casual, monologue) (low mumble) Uhm, and they work with images as well as text.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; no dominant emotion; style: casual, monologue; good recording, no background noise; genuineness 2.6/6; vocal-burst blend 3.2/10; 3.8s, EN.
EN_itLhZY5lb9y_W000072 · in -17.3 dBFS · gain -2.7 dB · emolia-02177
(casual, monologue) So for those who don't know, the Google, Google display ads can appear across over three million websites, over 650,000 apps, (ahem) um, and across Google properties such as Gmail and YouTube.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: casual, monologue; average recording, quiet background; genuineness 2.6/6; vocal-burst blend 1.7/10; 13.3s, EN.
EN_itLhZY5lb9y_W000073 · in -17.8 dBFS · gain -2.2 dB · emolia-02177
(interest, hope enthusiasm optimism, contentment · monologue, casual) They can also be super annoying on news sites, but they do actually work really well for not only brand awareness, but also for direct response when used in combination with remarketing audiences and a strong call to action. (low mumble) Um, and actually smart display campaigns is something I touch on later.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as interest, hope enthusiasm optimism, contentment; style: monologue, casual; average recording, quiet background; genuineness 2.7/6; vocal-burst blend 3.6/10; 18.1s, EN.
EN_itLhZY5lb9y_W000074 · in -17.2 dBFS · gain -2.8 dB · emolia-02177
Sexual Lust ↓  /  Reliefidentity +0.01 emotion 12 %   k-AB2-k3 · #17

This chain comes from the two-sided rule: it only counts if both emotions move — Sexual Lust down and Relief up — by at least 0.25 each.

The chain starts with Relief around average — 0.50, right about the corpus median — and ends with it strongly present at 0.87, higher than 87 % of clips in this corpus. That is a total rise of 0.37.

At the same time Sexual Lust goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.72 (higher than 72 % of clips in this corpus), a change of -0.27. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.22, then +0.15 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.73 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.81 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.73, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 25 s · en · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.709 before conversion and 0.724 after — it rose by 0.015. Neighbour-to-neighbour the worst pair went 0.663 → 0.626. (The earlier render, with segment 1 left raw, scores 0.632 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Relief moved +0.368 in the original and +0.044 after conversion — 12 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Sexual Lust, -0.270 became -0.051.

Quality. Mean predicted overall quality across the segments went 2.81 → 2.94 (+0.13) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.709 → 0.724 +0.015identity cos neighbours 0.663 → 0.626d_b rescored +0.368 → +0.044d_a rescored -0.270 → -0.051d_a mined -0.270d_b mined 0.369min_cos_consec (site) 0.8108min_cos_anchor (site) 0.7317dataset emolialang enspeaker EN_GDbv_jtEAsYtotal 23.9schain gain -0.6 dBseam step 0.4 dBcrossfades 100/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · good recording, no background noise, normal-paced, clear
(sexual lust, pain, contentment · very low-energy, relaxed, steady, ASMR) If you walk into a room with your head low and your words all jumbled, your insecurity can actually spread and might make others feel uncomfortable.
full caption & clip details
A young adult feminine voice; delivery is very low-energy, normal-paced, relaxed, steady; timbre is slightly warm, slightly bright, smooth, balanced body; clear, no disfluency, fairly narrow pitch, minimal breath; affect is mildly positive, slightly submissive, neutral openness; reads as sexual lust, pain, contentment; style: ASMR, whispered; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 1.4/10; 7.3s, EN.
EN_GDbv_jtEAsY_W000014 · in -17.1 dBFS · gain -3.0 dB · emolia-00695
(contemplation, concentration, sexual lust · very low-energy, slightly relaxed, fairly steady, whispered) Negativity, whether about yourself or others, tends to be perceived by others. We prefer positivity, just like how we appreciate people with self confidence. Number four. Have you stayed up to finish a paper or a deadline?
full caption & clip details
A young adult feminine voice; delivery is very low-energy, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; clear, little disfluency, fairly narrow pitch, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as contemplation, concentration, sexual lust; style: whispered, ASMR; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 2.3/10; 13.1s, EN.
EN_GDbv_jtEAsY_W000015 · in -17.7 dBFS · gain -2.3 dB · emolia-00695
(normally alert, slightly relaxed, fairly steady, casual) According to conducted research, you tend to look less attractive when you don't sleep.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, dark, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: casual, monologue; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 2.0/10; 3.9s, EN.
EN_GDbv_jtEAsY_W000016 · in -14.8 dBFS · gain -5.2 dB · emolia-00695
Infatuation ↓  /  Reliefidentity −0.02 emotion 43 %   k-AB2-k3 · #18

This chain comes from the two-sided rule: it only counts if both emotions move — Infatuation down and Relief up — by at least 0.25 each.

The chain starts with Relief around average — 0.50, right about the corpus median — and ends with it strongly present at 0.77, higher than 77 % of clips in this corpus. That is a total rise of 0.27.

At the same time Infatuation goes the other way, from 0.89 (higher than 89 % of clips in this corpus) to 0.51 (right about the corpus median), a change of -0.38. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.24, then +0.03 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.94 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.94 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.94), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 22 s · zh · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.915 before conversion and 0.891 after — it fell by 0.024. Neighbour-to-neighbour the worst pair went 0.915 → 0.891. (The earlier render, with segment 1 left raw, scores 0.908 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Relief moved +0.273 in the original and +0.119 after conversion — 43 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Infatuation, -0.378 became -0.374.

Quality. Mean predicted overall quality across the segments went 3.23 → 3.17 (-0.06) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.915 → 0.891 -0.024identity cos neighbours 0.915 → 0.891d_b rescored +0.273 → +0.119d_a rescored -0.378 → -0.374d_a mined -0.378d_b mined 0.273min_cos_consec (site) 0.9374min_cos_anchor (site) 0.9374dataset emolialang zhspeaker ZH_B00007_S02590total 21.1schain gain +2.1 dBseam step 1.1 dBcrossfades 150/100 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed
(normal-paced, formal, authoritative) 对外国商品的需求增加,使国内需求减少,从而导致美元贬值。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 4.4/10; 4.8s, ZH.
ZH_B00007_S02590_W000314 · in -25.9 dBFS · gain +5.9 dB · emolia-03342
(measured, authoritative, monologue) 第四,如果国内利率上升,从而向外国投资者提供更高收益,那么该国货币将在外汇市场上升值。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: authoritative, monologue; good recording, no background noise; genuineness 1.6/6; vocal-burst blend 4.7/10; 8.3s, ZH.
ZH_B00007_S02590_W000315 · in -25.9 dBFS · gain +5.9 dB · emolia-03342
(normal-paced, monologue, authoritative) 外国投资者为了收购海外公司而增加对美元的需求,从而供给更多他们国家的货币,以交换目标国家货币。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, authoritative; good recording, no background noise; genuineness 0.6/6; vocal-burst blend 4.8/10; 8.3s, ZH.
ZH_B00007_S02590_W000316 · in -25.7 dBFS · gain +5.7 dB · emolia-03342
Contemplation ↓  /  Hope Enthusiasm Optimismidentity −0.13 emotion 169 %   k-AB2-k3 · #19

This chain comes from the two-sided rule: it only counts if both emotions move — Contemplation down and Hope Enthusiasm Optimism up — by at least 0.25 each.

The chain starts with Hope Enthusiasm Optimism around average — 0.52, higher than 52 % of clips in this corpus — and ends with it strongly present at 0.87, higher than 87 % of clips in this corpus. That is a total rise of 0.36.

At the same time Contemplation goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.54 (higher than 54 % of clips in this corpus), a change of -0.44. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.16, then +0.20 — an even, steady climb — each clip carries about the same share.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.73 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.69 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.73, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 17 s · en · emolia

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.749 before conversion and 0.616 after — it fell by 0.133. Neighbour-to-neighbour the worst pair went 0.717 → 0.616. (The earlier render, with segment 1 left raw, scores 0.496 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Hope Enthusiasm Optimism moved +0.357 in the original and +0.602 after conversion — 169 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Contemplation, -0.438 became -0.811.

Quality. Mean predicted overall quality across the segments went 2.25 → 2.84 (+0.59) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.749 → 0.616 -0.133identity cos neighbours 0.717 → 0.616d_b rescored +0.357 → +0.602d_a rescored -0.438 → -0.811d_a mined -0.438d_b mined 0.357min_cos_consec (site) 0.6930min_cos_anchor (site) 0.7307dataset emolialang enspeaker EN_2gdr-7JW18ctotal 16.0schain gain +2.3 dBseam step 1.7 dBcrossfades 150/100 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, fairly smooth, balanced body, good recording, quiet background, normal-paced, normally alert, slightly relaxed
(contemplation, doubt · moderately variable, wide pitch range, casual, conversational) Given that these are all endogenous objects, perhaps you're thinking, what do I really mean when I say that?
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as contemplation, doubt; style: casual, conversational; good recording, quiet background; genuineness 2.8/6; vocal-burst blend 3.3/10; 5.8s, EN.
EN_2gdr-7JW18c_W000133 · in -15.0 dBFS · gain -5.0 dB · emolia-01907
(fairly steady, moderate pitch range, casual, monologue) And I think that's exactly a reasonable concern to have.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: casual, monologue; good recording, quiet background; genuineness 2.9/6; vocal-burst blend 2.3/10; 3.8s, EN.
EN_2gdr-7JW18c_W000136 · in -14.8 dBFS · gain -5.2 dB · emolia-01907
(fairly steady, moderate pitch range, conversational, playful) (ahem) And so let me do a little bit of an algebra that can allow us to come up with an equation where it's easier to see the.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; no dominant emotion; style: conversational, playful; good recording, quiet background; genuineness 2.0/6; vocal-burst blend 1.4/10; 6.8s, EN.
EN_2gdr-7JW18c_W000137 · in -14.3 dBFS · gain -5.7 dB · emolia-01907
Doubt ↓  /  Concentrationidentity +0.38 emotion 145 %   k-AB2-k3 · #20

This chain comes from the two-sided rule: it only counts if both emotions move — Doubt down and Concentration up — by at least 0.25 each.

The chain starts with Concentration clearly present — 0.62, higher than 62 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.31.

At the same time Doubt goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.69 (higher than 69 % of clips in this corpus), a change of -0.30. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.11, then +0.21 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores -0.11 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.06 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (-0.11, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 50 s · en · podcast

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was -0.099 before conversion and 0.278 after — it rose by 0.377. Neighbour-to-neighbour the worst pair went 0.113 → 0.390. (The earlier render, with segment 1 left raw, scores 0.143 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.312 in the original and +0.453 after conversion — 145 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Doubt, -0.300 became -0.397.

Quality. Mean predicted overall quality across the segments went 2.69 → 3.07 (+0.37) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 -0.099 → 0.278 +0.377identity cos neighbours 0.113 → 0.390d_b rescored +0.312 → +0.453d_a rescored -0.300 → -0.397d_a mined -0.300d_b mined 0.313min_cos_consec (site) 0.0594min_cos_anchor (site) -0.1135dataset podcastlang enspeaker 635275total 49.8schain gain +4.8 dBseam step 1.5 dBcrossfades 150/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a young adult feminine voice · neutral-toned, average recording, quiet background, frequent disfluency, moderate pitch range
(doubt, relief, embarrassment · measured, normally alert, relaxed, casual) Yeah. I I don't yet know feasibly I I don't yet know logistically how we would do that, but I'm okay with the concept so long as we could figure out how to enforce it.
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, relaxed, moderately variable; timbre is neutral-toned, dark, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is mildly negative, submissive, neutral openness; reads as doubt, relief, embarrassment; style: casual, conversational; average recording, quiet background; genuineness 5.1/6; vocal-burst blend 6.6/10; 11.2s, EN.
635275_00596848 · in -35.2 dBFS · gain +15.2 dB · podcast-02239
(embarrassment, doubt · normal-paced, normally alert, neutral tension, casual) Yeah, I mean this whole section repli is called replication encouragement. I I don't I don't think we want to encourage that, as Joan said.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is mildly negative, slightly submissive, neutral openness; reads as embarrassment, doubt; style: casual, conversational; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 2.7/10; 9.6s, EN.
635275_00612616 · in -34.3 dBFS · gain +14.3 dB · podcast-02236
(concentration · normal-paced, subdued, slightly relaxed, monologue) So we thought we were on safe ground because we were repeating language from the specific plan and from the zoning ordinance that talks about a maximum limit of 40% of the permitted FAR can be used for inclusionary commercial uses. And these are the commercial uses that meet the intents and the purposes of the specific plan. And it's it's basically supportive (low mumble) uh commercial uses in small amounts that serve the industrial uses and the waterfront uses.
full caption & clip details
A young adult masculine voice; delivery is subdued, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as concentration; style: monologue, casual; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 4.9/10; 29.3s, EN.
635275_00625704 · in -34.2 dBFS · gain +14.2 dB · podcast-03194