sc-AB2-k5 — voice-corrected

Two-sided AB2 after speaker cleaning (WavLM -id >= 0.80 on consecutive pairs AND against the first clip), k=5. This is the only source that ships min_cos_consec / min_cos_anchor.

This is the voice-conversion-corrected version of this tier. The original, uncorrected page is still there and unchanged: t_sc-AB2-k5.html. Every card below carries both renders so you can switch between them without leaving the page. What was done, and what it measured →
20chains converted
80segments re-voiced
0.825 → 0.817median worst-to-anchor identity cosine
85 %median emotion delta retained
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken.

(underlined, plain · delivery, style) — the tag before the words. Emotions first, then how it is delivered. Underlined descriptors are the ones that change across this chain — anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words are a real non-speech sound, printed where it happens.

The full generated caption for any clip is under “full caption & clip details”. Its perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.

The tags and captions describe the ORIGINAL clips — they are the corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the converted audio are the ones printed in each card's own paragraph and chips.
Concentration ↓  /  Astonishment Surpriseidentity −0.11 emotion REVERSED   sc-AB2-k5 · #1

This chain comes from the two-sided rule: it only counts if both emotions move — Concentration down and Astonishment Surprise up — by at least 0.25 each.

The chain starts with Astonishment Surprise below average — 0.36, lower than 64 % of clips in this corpus — and ends with it clearly present at 0.72, higher than 72 % of clips in this corpus. That is a total rise of 0.36.

At the same time Concentration goes the other way, from 0.90 (higher than 90 % of clips in this corpus) to 0.52 (higher than 52 % of clips in this corpus), a change of -0.39. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are -0.13, then +0.13, then +0.16, then +0.20 — not a clean run: step 1 moves back the other way by 0.13 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.88 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.85 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.88. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 39 s · zh · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.646 before conversion and 0.538 after — it fell by 0.107. Neighbour-to-neighbour the worst pair went 0.708 → 0.662. (The earlier render, with segment 1 left raw, scores 0.517 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Astonishment Surprise moved +0.364 in the original and +0.000 after conversion — 0 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Concentration, -0.387 became -0.524.

Quality. Mean predicted overall quality across the segments went 2.87 → 3.23 (+0.35) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.646 → 0.538 -0.107identity cos neighbours 0.708 → 0.662d_b rescored +0.364 → +0.000d_a rescored -0.387 → -0.524d_a mined -0.386d_b mined 0.364min_cos_consec (site) 0.8481min_cos_anchor (site) 0.8769dataset emolialang zhspeaker ZH_B00001_S02522total 37.9schain gain +2.2 dBseam step 1.9 dBcrossfades 150/150/150/100 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, no background noise, normally alert, slightly relaxed, no disfluency
(concentration · measured, fairly steady, monologue, formal) 一旦为人所知,机械中得以迅速的传播并淘汰了水中,但日晷保留了下来。作为检验新时钟效果的最后手段,早期的机械钟尚处在初期阶段,不精确且易于出故障。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: monologue, formal; average recording, no background noise; genuineness 1.0/6; vocal-burst blend 1.7/10; 19.8s, ZH.
ZH_B00001_S02522_W000063 · in -17.3 dBFS · gain -2.7 dB · emolia-03284
(measured, fairly steady, didactic, authoritative) 所以最好是在购买机械钟的同时购买一位智中工匠。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: didactic, authoritative; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 1.5/10; 5.2s, ZH.
ZH_B00001_S02522_W000064 · in -15.4 dBFS · gain -4.6 dB · emolia-03284
(measured, fairly steady, formal, monologue) 尽管在罗马帝国崩溃后,城市衰落的几个世纪里。
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 1.9/10; 5.0s, ZH.
ZH_B00001_S02522_W000065 · in -16.0 dBFS · gain -4.0 dB · emolia-03284
(normal-paced, fairly steady, formal, authoritative) 教会仪式使得人们保留了对纪实的兴趣。
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 1.4/6; vocal-burst blend 2.6/10; 3.8s, ZH.
ZH_B00001_S02522_W000066 · in -15.6 dBFS · gain -4.4 dB · emolia-03284
(measured, steady, formal, authoritative) 但教会时间是大自然的时间,白天黑夜不均分。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 1.1/6; vocal-burst blend 1.1/10; 5.0s, ZH.
ZH_B00001_S02522_W000067 · in -16.3 dBFS · gain -3.7 dB · emolia-03284
Contemplation ↓  /  Fatigue Exhaustionidentity −0.03 emotion 110 %   sc-AB2-k5 · #2

This chain comes from the two-sided rule: it only counts if both emotions move — Contemplation down and Fatigue Exhaustion up — by at least 0.25 each.

The chain starts with Fatigue Exhaustion around average — 0.55, higher than 55 % of clips in this corpus — and ends with it strongly present at 0.81, higher than 81 % of clips in this corpus. That is a total rise of 0.27.

At the same time Contemplation goes the other way, from 0.80 (higher than 80 % of clips in this corpus) to 0.51 (higher than 51 % of clips in this corpus), a change of -0.28. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.10, then +0.07, then +0.14, then -0.03 — not a clean run: step 4 moves back the other way by 0.03 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.90 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.89 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.90. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 32 s · zh · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.885 before conversion and 0.852 after — it fell by 0.033. Neighbour-to-neighbour the worst pair went 0.884 → 0.852. (The earlier render, with segment 1 left raw, scores 0.794 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fatigue Exhaustion moved +0.267 in the original and +0.293 after conversion — 110 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Contemplation, -0.284 became -0.299.

Quality. Mean predicted overall quality across the segments went 3.08 → 3.19 (+0.11) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.885 → 0.852 -0.033identity cos neighbours 0.884 → 0.852d_b rescored +0.267 → +0.293d_a rescored -0.284 → -0.299d_a mined -0.284d_b mined 0.267min_cos_consec (site) 0.8863min_cos_anchor (site) 0.8971dataset emolialang zhspeaker ZH_B00014_S07197total 30.9schain gain -0.3 dBseam step 0.7 dBcrossfades 150/150/150/100 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, no background noise, measured, normally alert, slightly relaxed, no disfluency
(fairly steady, formal, monologue) 总觉得周明畅的这个可能缺了点味道。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 1.2/6; vocal-burst blend 2.9/10; 3.7s, ZH.
ZH_B00014_S07197_W000043 · in -15.8 dBFS · gain -4.2 dB · emolia-03418
(confusion, doubt · steady, monologue, formal) 但已经不错了,好听还是好听的,只是苏晓自己对这首歌有感情歌,这矫情呢?
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as confusion, doubt; style: monologue, formal; average recording, no background noise; genuineness 0.9/6; vocal-burst blend 3.9/10; 7.5s, ZH.
ZH_B00014_S07197_W000044 · in -17.1 dBFS · gain -2.9 dB · emolia-03418
(fairly steady, narration, monologue) 放到这个世界的话,这个就是唯一的版本,相信大家会很喜欢的这首歌,不火就奇怪了。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, monologue; average recording, no background noise; genuineness 0.9/6; vocal-burst blend 4.2/10; 7.8s, ZH.
ZH_B00014_S07197_W000045 · in -16.4 dBFS · gain -3.6 dB · emolia-03418
(fairly steady, formal, narration) 苏小也提出了一些意见,让周明修改了一下,或许会更有感觉。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, narration; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 3.6/10; 5.6s, ZH.
ZH_B00014_S07197_W000046 · in -16.1 dBFS · gain -3.9 dB · emolia-03418
(fairly steady, formal, narration) 周明那边十分感激,觉得花这一千万还是挺划算的,苏小给他提供了不少帮助。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, narration; average recording, no background noise; genuineness 0.5/6; vocal-burst blend 3.5/10; 7.0s, ZH.
ZH_B00014_S07197_W000047 · in -16.6 dBFS · gain -3.4 dB · emolia-03418
Contemplation ↓  /  Concentrationidentity +0.05 emotion 107 %   sc-AB2-k5 · #3

This chain comes from the two-sided rule: it only counts if both emotions move — Contemplation down and Concentration up — by at least 0.25 each.

The chain starts with Concentration clearly present — 0.66, higher than 66 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.31.

At the same time Contemplation goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.68 (higher than 68 % of clips in this corpus), a change of -0.29. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.09, then +0.24, then -0.06, then +0.04 — not a clean run: step 3 moves back the other way by 0.06 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.86 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.88 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.86. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 51 s · de · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.766 before conversion and 0.820 after — it rose by 0.054. Neighbour-to-neighbour the worst pair went 0.815 → 0.849. (The earlier render, with segment 1 left raw, scores 0.726 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.307 in the original and +0.329 after conversion — 107 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Contemplation, -0.289 became -0.218.

Quality. Mean predicted overall quality across the segments went 2.84 → 3.11 (+0.27) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.766 → 0.820 +0.054identity cos neighbours 0.815 → 0.849d_b rescored +0.307 → +0.329d_a rescored -0.289 → -0.218d_a mined -0.290d_b mined 0.308min_cos_consec (site) 0.8780min_cos_anchor (site) 0.8605dataset emolialang despeaker DE_QVtYzDjCWZ0total 49.5schain gain +3.0 dBseam step 3.1 dBcrossfades 150/150/150/100 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a middle-aged masculine voice · neutral-toned, balanced body, normally alert, slightly relaxed, moderate pitch range
(contemplation, confusion, triumph · measured, fairly steady, frequent disfluency, storytelling) Und über diesen zeitlichen Fluss habe ich mir letztlich unsere Wörter geprägt. Sprich, ich habe also aus der Analogie
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation, confusion, triumph; style: storytelling, monologue; average recording, quiet background; genuineness 2.8/6; vocal-burst blend 0.0/10; 7.8s, DE.
DE_QVtYzDjCWZ0_W000072 · in -18.4 dBFS · gain -1.6 dB · emolia-00255
(pride, malevolence malice, bitterness · measured, moderately variable, some disfluency, monologue) Unser einzelne Wörter mit meinem alltäglichen Wahnsinn, den ich fast jeden Morgen erlebe, ein mentales Modell konstruiert, ein mentales Modell generiert.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as pride, malevolence malice, bitterness; style: monologue, didactic; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 0.0/10; 9.4s, DE.
DE_QVtYzDjCWZ0_W000073 · in -17.2 dBFS · gain -2.8 dB · emolia-00255
(concentration, emotional numbness, sourness · measured, fairly steady, some disfluency, didactic) Das ich dann sozusagen in mein Langzeitgedächtnis, in mein Langzeitgedächtnis vorhandenen und abgespeicherten Erfahrungs- und Wissensstrukturen eingeordnet und angedockt habe. Das heißt, in diesem Sinne ist eigentlich Lernen ein aktiver, ein konstruktiver Prozess.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, emotional numbness, sourness; style: didactic, monologue; average recording, no background noise; genuineness 2.6/6; vocal-burst blend 0.0/10; 17.7s, DE.
DE_QVtYzDjCWZ0_W000074 · in -19.6 dBFS · gain -0.4 dB · emolia-00255
(concentration · measured, fairly steady, frequent disfluency, didactic) Bei dem wir letztlich relevante Informationen selektieren, diese Informationen organisieren.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: didactic, monologue; average recording, quiet background; genuineness 2.3/6; vocal-burst blend 0.0/10; 6.2s, DE.
DE_QVtYzDjCWZ0_W000075 · in -17.8 dBFS · gain -2.2 dB · emolia-00255
(concentration · normal-paced, fairly steady, little disfluency, didactic) Und letztlich diese organisierten Informationen in unsere bereits vorhandenen Wissens- und Erfahrungsstrukturen unseres Langzeitgedächtnisses integrieren.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as concentration; style: didactic, monologue; good recording, no background noise; genuineness 1.4/6; vocal-burst blend 0.0/10; 9.2s, DE.
DE_QVtYzDjCWZ0_W000076 · in -19.6 dBFS · gain -0.4 dB · emolia-00255
Confusion ↓  /  Concentrationidentity +0.03 emotion 75 %   sc-AB2-k5 · #4

This chain comes from the two-sided rule: it only counts if both emotions move — Confusion down and Concentration up — by at least 0.25 each.

The chain starts with Concentration clearly present — 0.62, higher than 62 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.35.

At the same time Confusion goes the other way, from 0.87 (higher than 87 % of clips in this corpus) to 0.62 (higher than 62 % of clips in this corpus), a change of -0.25. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.16, then -0.01, then +0.08, then +0.12 — not a clean run: step 2 moves back the other way by 0.01 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.94 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.94 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.94), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 54 s · zh · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.910 before conversion and 0.936 after — it rose by 0.026. Neighbour-to-neighbour the worst pair went 0.920 → 0.910. (The earlier render, with segment 1 left raw, scores 0.852 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.352 in the original and +0.265 after conversion — 75 % of the delta retained, which is most of it. On the other named axis, Confusion, -0.251 became -0.815.

Quality. Mean predicted overall quality across the segments went 3.20 → 3.30 (+0.10) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.910 → 0.936 +0.026identity cos neighbours 0.920 → 0.910d_b rescored +0.352 → +0.265d_a rescored -0.251 → -0.815d_a mined -0.250d_b mined 0.351min_cos_consec (site) 0.9416min_cos_anchor (site) 0.9435dataset emolialang zhspeaker ZH_B00017_S04695total 52.3schain gain +2.3 dBseam step 1.2 dBcrossfades 100/100/100/100 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, no background noise, normal-paced, normally alert, slightly relaxed
(monologue, formal) 第十章讨论蓝海战略的更新及动态特性,这涉及业务层面以及拥有多业务企业的公司层面。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 4.2/10; 7.5s, ZH.
ZH_B00017_S04695_W000093 · in -16.7 dBFS · gain -3.3 dB · emolia-03447
(monologue, formal) 在扩展版中,我们将原有的讨论扩展到在纵向时间段里,如何管理和监测你的个别业务,以及你的公司业务组合,以令企业持续保持上乘的业绩。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 4.5/10; 11.6s, ZH.
ZH_B00017_S04695_W000094 · in -18.6 dBFS · gain -1.4 dB · emolia-03447
(monologue, formal) 这一章解决的是有关战略更新的风险管理这一重要问题,旨在令追求蓝海战略的过程制度化,而不是昙花一现。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 2.6/10; 9.0s, ZH.
ZH_B00017_S04695_W000095 · in -18.9 dBFS · gain -1.1 dB · emolia-03447
(monologue, formal) 这一章展示了纵向时间段下在管理企业业务组合的过程中,红海战略和蓝海战略如何能够相互契合和相互补充。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; average recording, no background noise; genuineness 0.8/6; vocal-burst blend 3.3/10; 9.6s, ZH.
ZH_B00017_S04695_W000096 · in -18.1 dBFS · gain -1.9 dB · emolia-03447
(concentration · monologue, formal) 最后,我们新添一章为扩展版作结。该章详细讨论了十个最常见的红海陷阱,企业组织起航,驶向蓝海之际,常因这些陷阱而无法脱离红海。在此,我们明确指出,如何避免落入这些陷阱。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: monologue, formal; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 4.9/10; 15.3s, ZH.
ZH_B00017_S04695_W000097 · in -18.8 dBFS · gain -1.2 dB · emolia-03447
Interest ↓  /  Shameidentity −0.02 emotion 69 %   sc-AB2-k5 · #5

This chain comes from the two-sided rule: it only counts if both emotions move — Interest down and Shame up — by at least 0.25 each.

The chain starts with Shame clearly present — 0.68, higher than 68 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.29.

At the same time Interest goes the other way, from 0.91 (higher than 91 % of clips in this corpus) to 0.60 (higher than 60 % of clips in this corpus), a change of -0.31. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.17, then +0.10, then +0.02, then -0.01 — not a clean run: step 4 moves back the other way by 0.01 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.96 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.97 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.96), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 81 s · sr · eurospeech

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.960 before conversion and 0.937 after — it fell by 0.023. Neighbour-to-neighbour the worst pair went 0.962 → 0.930. (The earlier render, with segment 1 left raw, scores 0.820 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Shame moved +0.285 in the original and +0.196 after conversion — 69 % of the delta retained. On the other named axis, Interest, -0.311 became -0.130.

Quality. Mean predicted overall quality across the segments went 3.16 → 3.24 (+0.08) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.960 → 0.937 -0.023identity cos neighbours 0.962 → 0.930d_b rescored +0.285 → +0.196d_a rescored -0.311 → -0.130d_a mined -0.311d_b mined 0.285min_cos_consec (site) 0.9664min_cos_anchor (site) 0.9585dataset eurospeechlang srspeaker serbia_serbia_2020_290_120total 80.1schain gain +3.0 dBseam step 1.7 dBcrossfades 150/100/150/150 ms
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: a child masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, quiet background, normally alert, slightly relaxed
(interest · fast, authoritative, monologue) krivičnom gonjenju i sprečavanju kriminala kroz saradnju i uzajamnu pravnu pomoć u krivičnim stvarima i na taj način doprineti efikasnoj borbi protiv svih vidova kriminala i (ahem) omogućiti efikasnije procesuiranje počinilaca krivičnih dela.Prema
full caption & clip details
A child masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as interest; style: authoritative, monologue; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 5.1/10; 15.2s, SR.
serbia_serbia_2020_290_12052021_911856_927088 · in -18.0 dBFS · gain -2.0 dB · eurospeech-02766
(thankfulness gratitude, disappointment · normal-paced, monologue, authoritative) dela.Prema navedenom Ugovoru, strane će u skladu sa odredbama Ugovora, pružiti jedna drugoj najširu moguću pravnu pomoć u krivičnim stvarima.Obim pravne pomoći koja je predviđena ugovorom obuhvata uručenje dokumenata, pronalaženje, identifikovanje lica i predmeta, pribavljanje informacija, dokumenata i dokaza, obezbeđivanje
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as thankfulness gratitude, disappointment; style: monologue, authoritative; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 4.9/10; 20.0s, SR.
serbia_serbia_2020_290_12052021_927088_947088 · in -18.7 dBFS · gain -1.3 dB · eurospeech-02766
(relief, shame, thankfulness gratitude · fast, monologue, authoritative) predmeta, pretres i konfiskaciju, izvođenje dokaza i pribavljanje izjava, video-konferencijsko saslušanje, pronalaženje, oduzimanje i zaplenu prihoda od kriminala i bilo koji drugi oblik pomoći koji nije u suprotnosti sa zakonom zemlje molilje.Ugovor
full caption & clip details
A child masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as relief, shame, thankfulness gratitude; style: monologue, authoritative; average recording, quiet background; genuineness 2.1/6; vocal-burst blend 4.2/10; 16.3s, SR.
serbia_serbia_2020_290_12052021_947088_963424 · in -21.6 dBFS · gain +1.6 dB · eurospeech-02766
(pride, shame, triumph · fast, authoritative, didactic) sadrži i odredbe koje uređuju način izvršenja zamolnica, formu i sadržaj zamolnica, ograničenje pružanja pravne pomoći, pronalaženje, identifikovanje lica i predmeta, uručenje sudskih pismena i dokumenata,
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as pride, shame, triumph; style: authoritative, didactic; average recording, quiet background; genuineness 2.7/6; vocal-burst blend 5.4/10; 13.5s, SR.
serbia_serbia_2020_290_12052021_963424_976944 · in -17.3 dBFS · gain -2.7 dB · eurospeech-02766
(shame, thankfulness gratitude, pride · fast, monologue, authoritative) pribavljanje informacija, dokumenata i predmeta, zaplenu predmeta, uzimanje izjava od zamoljene strane, predaju pritvorenih lica radi svedočenja strani molilji, davanje izjava strani molilji, saslušanje putem video-konferencijske veze,
full caption & clip details
A young adult masculine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as shame, thankfulness gratitude, pride; style: monologue, authoritative; average recording, quiet background; genuineness 1.7/6; vocal-burst blend 4.6/10; 15.8s, SR.
serbia_serbia_2020_290_12052021_976944_992736 · in -17.9 dBFS · gain -2.1 dB · eurospeech-02766
Emotional Numbness ↓  /  Infatuationidentity −0.05 emotion 154 %   sc-AB2-k5 · #6

This chain comes from the two-sided rule: it only counts if both emotions move — Emotional Numbness down and Infatuation up — by at least 0.25 each.

The chain starts with Infatuation clearly present — 0.62, higher than 62 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.35.

At the same time Emotional Numbness goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.71 (higher than 71 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.11, then +0.22, then -0.01, then +0.02 — not a clean run: step 3 moves back the other way by 0.01 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.82 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.89 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.82. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 70 s · de · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.642 before conversion and 0.591 after — it fell by 0.051. Neighbour-to-neighbour the worst pair went 0.642 → 0.591. (The earlier render, with segment 1 left raw, scores 0.662 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Infatuation moved +0.347 in the original and +0.535 after conversion — 154 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Emotional Numbness, -0.265 became +0.017.

Quality. Mean predicted overall quality across the segments went 2.90 → 3.17 (+0.27) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.642 → 0.591 -0.051identity cos neighbours 0.642 → 0.591d_b rescored +0.347 → +0.535d_a rescored -0.265 → +0.017d_a mined -0.265d_b mined 0.347min_cos_consec (site) 0.8911min_cos_anchor (site) 0.8186dataset emolialang despeaker DE_B00001_S00064total 69.0schain gain +1.4 dBseam step 1.9 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult feminine voice · neutral-toned, neutral-bright, slightly relaxed, fairly steady, clear, light breath
(emotional numbness, disappointment, fear · measured, normally alert, no disfluency, ASMR) Seit der Krieg zu Ende war, hatte sie sich mit allen möglichen Jobs durchgeschlagen.
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, disappointment, fear; style: ASMR, narration; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 0.0/10; 5.3s, DE.
DE_B00001_S00064_W000021 · in -22.1 dBFS · gain +2.0 dB · emolia-00044
(awe, longing, concentration · measured, very low-energy, no disfluency, whispered) An ihrem Beruf als Straßenbahn-Schaffnerin, den sie seit ein paar Jahren hatte, mochte sie die Uniformen und die Bewegung, den Wechsel der Bilder und das Rollen unter den Füßen. Sonst mochte sie ihn nicht. Sie hatte keine Familie. Sie war 36.
full caption & clip details
An elderly feminine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is mildly positive, neutral stance, neutral openness; reads as awe, longing, concentration; style: whispered, narration; average recording, quiet background; genuineness 0.4/6; vocal-burst blend 0.0/10; 19.3s, DE.
DE_B00001_S00064_W000022 · in -22.2 dBFS · gain +2.2 dB · emolia-00044
(contempt, longing, sourness · measured, normally alert, no disfluency, whispered) Das alles entzählte sie, als sei es nicht ihr Leben, sondern das Leben eines anderen, den sie nicht gut kennt und der sie nicht angeht.
full caption & clip details
A child feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as contempt, longing, sourness; style: whispered, ASMR; average recording, quiet background; genuineness 0.5/6; vocal-burst blend 0.0/10; 9.4s, DE.
DE_B00001_S00064_W000023 · in -22.1 dBFS · gain +2.1 dB · emolia-00044
(fear, malevolence malice, disappointment · measured, very low-energy, some disfluency, narration) Was ich genauer wissen wollte, wusste sie oft nicht mehr und sie verstand auch nicht, warum mich interessierte, was aus ihren Eltern geworden war, ob sie Geschwister gehabt, wie sie in Berlin gelebt und was sie bei den Soldaten gemacht hatte. Was du alles wissen willst, Junchen. Ebenso war es mit der Zukunft.
full caption & clip details
An elderly feminine voice; delivery is very low-energy, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; clear, some disfluency, fairly narrow pitch, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as fear, malevolence malice, disappointment; style: narration, whispered; average recording, quiet background; genuineness 1.5/6; vocal-burst blend 0.9/10; 23.4s, DE.
DE_B00001_S00064_W000024 · in -22.7 dBFS · gain +2.7 dB · emolia-00044
(infatuation, sexual lust, contentment · normal-paced, normally alert, no disfluency, narration) Natürlich schmiedete ich keine Pläne für Heirat und Familie, aber ich nahm eine Beziehung von Julianne Sorelle zu Madame de Renal mehr an Teil als an der zu mithilfe der Maud.
full caption & clip details
An elderly feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as infatuation, sexual lust, contentment; style: narration, whispered; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 0.0/10; 12.3s, DE.
DE_B00001_S00064_W000025 · in -20.6 dBFS · gain +0.6 dB · emolia-00044
Fatigue Exhaustion ↓  /  Concentrationidentity −0.02 emotion 106 %   sc-AB2-k5 · #7

This chain comes from the two-sided rule: it only counts if both emotions move — Fatigue Exhaustion down and Concentration up — by at least 0.25 each.

The chain starts with Concentration clearly present — 0.62, higher than 62 % of clips in this corpus — and ends with it strongly present at 0.87, higher than 87 % of clips in this corpus. That is a total rise of 0.25.

At the same time Fatigue Exhaustion goes the other way, from 0.81 (higher than 81 % of clips in this corpus) to 0.35 (lower than 65 % of clips in this corpus), a change of -0.46. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.08, then -0.16, then +0.09, then +0.24 — not a clean run: step 2 moves back the other way by 0.16 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.81 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.84 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.81. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 36 s · zh · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.764 before conversion and 0.748 after — it fell by 0.016. Neighbour-to-neighbour the worst pair went 0.767 → 0.745. (The earlier render, with segment 1 left raw, scores 0.685 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.254 in the original and +0.270 after conversion — 106 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Fatigue Exhaustion, -0.460 became -0.571.

Quality. Mean predicted overall quality across the segments went 3.18 → 3.23 (+0.06) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.764 → 0.748 -0.016identity cos neighbours 0.767 → 0.745d_b rescored +0.254 → +0.270d_a rescored -0.460 → -0.571d_a mined -0.460d_b mined 0.254min_cos_consec (site) 0.8405min_cos_anchor (site) 0.8078dataset emolialang zhspeaker ZH_B00008_S02941total 34.3schain gain +0.6 dBseam step 1.2 dBcrossfades 100/100/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, no background noise, measured, normally alert, slightly relaxed
(fairly steady, little disfluency, clear, monologue) 按照世界银行的口径,二零二一年中国人均gdp达到一万两千五百五十六美元。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; good recording, no background noise; genuineness 1.4/6; vocal-burst blend 3.4/10; 5.9s, ZH.
ZH_B00008_S02941_W000018 · in -25.1 dBFS · gain +5.1 dB · emolia-03353
(fairly steady, no disfluency, clear, monologue) 已经接近于高收入国家门槛水平。大致体来说,我们可以把人均GDP处于一万两千到两万三千美元区间的国家看作。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, didactic; average recording, no background noise; genuineness 1.3/6; vocal-burst blend 3.0/10; 8.8s, ZH.
ZH_B00008_S02941_W000019 · in -25.9 dBFS · gain +5.9 dB · emolia-03353
(emotional numbness, pain · steady, no disfluency, clear, formal) 对高收入国家进行三等分的第一组,别。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, pain; style: formal, monologue; very good recording, no background noise; genuineness 0.6/6; vocal-burst blend 1.9/10; 3.3s, ZH.
ZH_B00008_S02941_W000020 · in -26.6 dBFS · gain +6.6 dB · emolia-03353
(fairly steady, some disfluency, somewhat unclear, monologue) 从这组别的门槛水平起步,再到达第二个三等分组的门槛水平,意味着进入中等发达国家行列。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, didactic; average recording, no background noise; genuineness 1.6/6; vocal-burst blend 2.6/10; 8.2s, ZH.
ZH_B00008_S02941_W000021 · in -26.7 dBFS · gain +6.7 dB · emolia-03353
(fairly steady, little disfluency, clear, monologue) 这是中国在二零三五年要实现的远景目标,因此可以形象的把这个发展阶段成为现代化从门槛到中途的阶段。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, didactic; average recording, no background noise; genuineness 1.3/6; vocal-burst blend 4.9/10; 8.7s, ZH.
ZH_B00008_S02941_W000022 · in -24.7 dBFS · gain +4.7 dB · emolia-03353
Contemplation ↓  /  Fatigue Exhaustionidentity −0.03 emotion 102 %   sc-AB2-k5 · #8

This chain comes from the two-sided rule: it only counts if both emotions move — Contemplation down and Fatigue Exhaustion up — by at least 0.25 each.

The chain starts with Fatigue Exhaustion clearly present — 0.62, higher than 62 % of clips in this corpus — and ends with it at the very top of the corpus at 0.91, higher than 91 % of clips in this corpus. That is a total rise of 0.29.

At the same time Contemplation goes the other way, from 0.90 (higher than 90 % of clips in this corpus) to 0.52 (higher than 52 % of clips in this corpus), a change of -0.38. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.05, then +0.18, then -0.01, then +0.07 — not a clean run: step 3 moves back the other way by 0.01 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.93 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.93 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.93), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 51 s · zh · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.930 before conversion and 0.899 after — it fell by 0.031. Neighbour-to-neighbour the worst pair went 0.937 → 0.897. (The earlier render, with segment 1 left raw, scores 0.833 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fatigue Exhaustion moved +0.287 in the original and +0.292 after conversion — 102 % of the delta retained, which is essentially all of it. On the other named axis, Contemplation, -0.385 became -0.359.

Quality. Mean predicted overall quality across the segments went 3.13 → 3.23 (+0.10) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.930 → 0.899 -0.031identity cos neighbours 0.937 → 0.897d_b rescored +0.287 → +0.292d_a rescored -0.385 → -0.359d_a mined -0.386d_b mined 0.287min_cos_consec (site) 0.9315min_cos_anchor (site) 0.9275dataset emolialang zhspeaker ZH_B00064_S07480total 49.1schain gain +1.2 dBseam step 1.0 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, slightly dark, fairly smooth, average recording, no background noise, measured, normally alert, slightly relaxed
(contemplation · fairly steady, some disfluency, clear, monologue) 当他再一次回到轮回中枢后,他能够感觉到自己已经能够原地跳起三米左右,高度超越了人类极限。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation; style: monologue, narration; average recording, no background noise; genuineness 1.2/6; vocal-burst blend 4.4/10; 10.2s, ZH.
ZH_B00064_S07480_W000034 · in -20.6 dBFS · gain +0.6 dB · emolia-03912
(concentration, infatuation · steady, some disfluency, somewhat unclear, monologue) 同时,速度、力量皆是大为提升。按照龙辉世界的实力评价,只是身体素质这方面,他就已经突破了f达到了一级。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, infatuation; style: monologue, didactic; average recording, no background noise; genuineness 1.2/6; vocal-burst blend 3.3/10; 12.6s, ZH.
ZH_B00064_S07480_W000035 · in -21.6 dBFS · gain +1.6 dB · emolia-03912
(sexual lust, intoxication altered states of consciousness · steady, no disfluency, clear, monologue) 新获得的念动力影响范围并不强,并不能隔空杀人,仅有五米左右,立刻致人死亡的作用。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as sexual lust, intoxication altered states of consciousness; style: monologue, whispered; average recording, no background noise; genuineness 0.8/6; vocal-burst blend 3.5/10; 9.5s, ZH.
ZH_B00064_S07480_W000036 · in -20.5 dBFS · gain +0.5 dB · emolia-03912
(concentration · steady, no disfluency, clear, monologue) 十米之内能够稍微操控一下几斤重量的物品,能够使用三次,目前来说作用不小。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration; style: monologue, whispered; average recording, no background noise; genuineness 1.1/6; vocal-burst blend 3.0/10; 9.6s, ZH.
ZH_B00064_S07480_W000037 · in -20.6 dBFS · gain +0.6 dB · emolia-03912
(fatigue exhaustion, confusion · fairly steady, no disfluency, clear, monologue) 用来出其不意,取得敌人的头发、血液,或者悄悄杀死敌人,还是极为有用的。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, thin; clear, no disfluency, moderate pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; reads as fatigue exhaustion, confusion; style: monologue, narration; average recording, no background noise; genuineness 1.1/6; vocal-burst blend 3.8/10; 8.1s, ZH.
ZH_B00064_S07480_W000038 · in -21.4 dBFS · gain +1.4 dB · emolia-03912
Relief ↓  /  Shameidentity −0.00 emotion 88 %   sc-AB2-k5 · #9

This chain comes from the two-sided rule: it only counts if both emotions move — Relief down and Shame up — by at least 0.25 each.

The chain starts with Shame clearly present — 0.72, higher than 72 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.26.

At the same time Relief goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.71 (higher than 71 % of clips in this corpus), a change of -0.27. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.17, then +0.06, then -0.06, then +0.09 — not a clean run: step 3 moves back the other way by 0.06 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.92 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.92 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.92), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 105 s · pt · podcast

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.935 before conversion and 0.931 after — it fell by 0.005. Neighbour-to-neighbour the worst pair went 0.939 → 0.931. (The earlier render, with segment 1 left raw, scores 0.850 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Shame moved +0.261 in the original and +0.230 after conversion — 88 % of the delta retained, which is most of it. On the other named axis, Relief, -0.269 became -0.467.

Quality. Mean predicted overall quality across the segments went 3.01 → 3.27 (+0.26) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.935 → 0.931 -0.005identity cos neighbours 0.939 → 0.931d_b rescored +0.261 → +0.230d_a rescored -0.269 → -0.467d_a mined -0.269d_b mined 0.260min_cos_consec (site) 0.9212min_cos_anchor (site) 0.9212dataset podcastlang ptspeaker 887678total 103.7schain gain +4.1 dBseam step 1.4 dBcrossfades 150/100/150/150 ms
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · balanced body, average recording, quiet background, measured, frequent disfluency, somewhat unclear
(relief, triumph · normally alert, slightly relaxed, fairly steady, monologue) havia essa aspectativa. Ela se confirmou. O esporte entrou com três zagueiros. Mas aí a continuação da formação conservadora seria a tirada, a saída de Patrick por problemas que a gente já sabe. Foi punido com o terceiro cartão amarelo. Então ele estava suspenso.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as relief, triumph; style: monologue, playful; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 3.8/10; 21.1s, PT.
887678_00034356 · in -28.9 dBFS · gain +8.9 dB · podcast-03279
(disgust, interest · normally alert, slightly relaxed, moderately variable, cartoonish) time. Porque Raul Prata tem um sistema defensivo mais sólido do que Patrick. Aliás, eu já venho falando aqui que Patrick está tendo uma dificuldade de apresentar um futebol (low mumble)
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as disgust, interest; style: cartoonish, didactic; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 1.1/10; 15.9s, PT.
887678_00037724 · in -28.5 dBFS · gain +8.5 dB · podcast-02725
(disappointment, anger, sourness · normally alert, slightly relaxed, moderately variable, cartoonish) (low mumble) mais sólido no esporte. Talvez essas últimas 10 partidas que o esporte deu, um pouco menos ou um pouco mais, enfim. Mas Patrick tem demonstrado diretamente essa dificuldade. Ele tem caído de produção. Chegou
full caption & clip details
An elderly masculine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is slightly cool, slightly dark, rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, slightly dominant, slightly guarded; reads as disappointment, anger, sourness; style: cartoonish, didactic; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 1.8/10; 19.0s, PT.
887678_00039312 · in -28.9 dBFS · gain +8.9 dB · podcast-02744
(contentment, relief, disappointment · normally alert, neutral tension, moderately variable, cartoonish) (low mumble) Tecnicamente ele vem apresentado abaixo, no meu ponto de vista, do que Raul Prata. Mas ele não tirou Patrick por Patrick ser uma voz de liderança muito forte no grupo. (low mumble) Está na reta final. O próprio Jair Ventura não quer perder esse grupo. O grupo está fechado, unido. E
full caption & clip details
A child masculine voice; delivery is normally alert, measured, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as contentment, relief, disappointment; style: cartoonish, playful; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 4.6/10; 22.1s, PT.
887678_00043580 · in -28.4 dBFS · gain +8.4 dB · podcast-02720
(shame, contentment, bitterness · subdued, slightly relaxed, fairly steady, monologue) isso poderia mexer um pouco com a estrutura psicológica do time. Então Jair resolveu abraçar essa ideia, pagar esse preço das baixas apresentações que Patrick vem dando no esporte. Mas enfim, conservadoramente falando, poderia entrar com o Raul Prata. Raul Prata sustentaria, dava uma base melhor.
full caption & clip details
A child masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as shame, contentment, bitterness; style: monologue; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 3.3/10; 26.3s, PT.
887678_00045784 · in -27.5 dBFS · gain +7.5 dB · podcast-02739
Contemplation ↓  /  Fatigue Exhaustionidentity −0.02 emotion 65 %   sc-AB2-k5 · #10

This chain comes from the two-sided rule: it only counts if both emotions move — Contemplation down and Fatigue Exhaustion up — by at least 0.25 each.

The chain starts with Fatigue Exhaustion around average — 0.47, lower than 53 % of clips in this corpus — and ends with it strongly present at 0.90, higher than 90 % of clips in this corpus. That is a total rise of 0.42.

At the same time Contemplation goes the other way, from 0.82 (higher than 82 % of clips in this corpus) to 0.48 (lower than 52 % of clips in this corpus), a change of -0.34. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.15, then +0.21, then -0.14, then +0.21 — not a clean run: step 3 moves back the other way by 0.14 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.90 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.88 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.90), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 30 s · zh · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.847 before conversion and 0.832 after — it fell by 0.015. Neighbour-to-neighbour the worst pair went 0.826 → 0.791. (The earlier render, with segment 1 left raw, scores 0.718 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fatigue Exhaustion moved +0.399 in the original and +0.260 after conversion — 65 % of the delta retained. On the other named axis, Contemplation, -0.341 became -0.025.

Quality. Mean predicted overall quality across the segments went 3.10 → 3.15 (+0.05) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.847 → 0.832 -0.015identity cos neighbours 0.826 → 0.791d_b rescored +0.399 → +0.260d_a rescored -0.341 → -0.025d_a mined -0.340d_b mined 0.422min_cos_consec (site) 0.8782min_cos_anchor (site) 0.9004dataset emolialang zhspeaker ZH_B00011_S06613total 28.5schain gain +0.2 dBseam step 0.7 dBcrossfades 100/150/100/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed, fairly steady
(measured, narration, monologue) 长老们各持理由不断,辩论,争吵着最后一个勉强,算是好的消息传来。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, monologue; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 2.6/10; 7.9s, ZH.
ZH_B00011_S06613_W000034 · in -21.4 dBFS · gain +1.4 dB · emolia-03384
(measured, narration, monologue) 这是伊莱雅斯风尘仆仆赶回马拉卡金亲自传回的消息,一个月的连续高强度作战。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, monologue; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 2.5/10; 9.0s, ZH.
ZH_B00011_S06613_W000035 · in -20.8 dBFS · gain +0.8 dB · emolia-03384
(measured, narration, formal) 这个强悍的牛头人战士又多了几处伤疤。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, formal; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 3.0/10; 4.1s, ZH.
ZH_B00011_S06613_W000036 · in -20.0 dBFS · gain -0.0 dB · emolia-03384
(confusion · normal-paced, narration, formal) 根据前线的猛禽,德鲁伊们持续不断的空中侦查。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as confusion; style: narration, formal; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 4.6/10; 4.1s, ZH.
ZH_B00011_S06613_W000037 · in -21.3 dBFS · gain +1.3 dB · emolia-03384
(normal-paced, narration, formal) 鹰身人和牛头人一样面临着巨大的后勤压力。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: narration, formal; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 5.0/10; 4.0s, ZH.
ZH_B00011_S06613_W000038 · in -19.7 dBFS · gain -0.3 dB · emolia-03384
Hope Enthusiasm Optimism ↓  /  Fatigue Exhaustionidentity −0.06 emotion 85 %   sc-AB2-k5 · #11

This chain comes from the two-sided rule: it only counts if both emotions move — Hope Enthusiasm Optimism down and Fatigue Exhaustion up — by at least 0.25 each.

The chain starts with Fatigue Exhaustion clearly present — 0.67, higher than 67 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.32.

At the same time Hope Enthusiasm Optimism goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.73 (higher than 73 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are -0.07, then +0.12, then +0.24, then +0.04 — not a clean run: step 1 moves back the other way by 0.07 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.89 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.88 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.89. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 61 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.851 before conversion and 0.788 after — it fell by 0.063. Neighbour-to-neighbour the worst pair went 0.851 → 0.782. (The earlier render, with segment 1 left raw, scores 0.617 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Fatigue Exhaustion moved +0.324 in the original and +0.277 after conversion — 85 % of the delta retained, which is most of it. On the other named axis, Hope Enthusiasm Optimism, -0.256 became -0.133.

Quality. Mean predicted overall quality across the segments went 2.91 → 3.14 (+0.24) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.851 → 0.788 -0.063identity cos neighbours 0.851 → 0.782d_b rescored +0.324 → +0.277d_a rescored -0.256 → -0.133d_a mined -0.256d_b mined 0.324min_cos_consec (site) 0.8776min_cos_anchor (site) 0.8918dataset emolialang enspeaker EN_5nsC7k4-UwEtotal 59.3schain gain +2.5 dBseam step 1.0 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 4 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, balanced body, average recording, quiet background, fairly steady, moderate pitch range, light breath
(hope enthusiasm optimism, interest, contentment · normal-paced, normally alert, slightly relaxed, casual) So that's what you guys are gonna see. A couple of videos with low membership are our low people on server. It's not gonna be the normal 20 to 30 people, (low mumble) uhm, plus that we usually get, but it's a start. It's somewhere that we can start and expand and grow from.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as hope enthusiasm optimism, interest, contentment; style: casual, monologue; average recording, quiet background; genuineness 3.3/6; vocal-burst blend 5.3/10; 14.7s, EN.
EN_5nsC7k4-UwE_W000008 · in -18.3 dBFS · gain -1.7 dB · emolia-00998
(elation, hope enthusiasm optimism, malevolence malice · brisk, normally alert, slightly relaxed, casual) So if we can start patrols going from 12 to 12, that'll be amazing. And then people can just hop on and get off whenever they like during the day.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as elation, hope enthusiasm optimism, malevolence malice; style: casual, monologue; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 5.4/10; 7.0s, EN.
EN_5nsC7k4-UwE_W000009 · in -18.5 dBFS · gain -1.5 dB · emolia-00998
(embarrassment, hope enthusiasm optimism · normal-paced, normally alert, neutral tension, casual) But (low mumble) uhm, yeaah, that's what we're gonna be trying to do now. Remember to follow me on my socials, link in the description down below. If you're in another time zone, don't worry, you can still join the VRP.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as embarrassment, hope enthusiasm optimism; style: casual, conversational; average recording, quiet background; genuineness 4.6/6; vocal-burst blend 8.3/10; 8.6s, EN.
EN_5nsC7k4-UwE_W000010 · in -18.7 dBFS · gain -1.3 dB · emolia-00998
(fatigue exhaustion, affection, longing · normal-paced, normally alert, relaxed, casual) (low mumble) Uhm, most of the content creators would either be coming on, (low mumble) uhm, earlier, cause we would actually like to come on earlier, but, you know, it is what it is. (low mumble) Uhm, we just need some of you guys to come on and
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, slightly guarded; reads as fatigue exhaustion, affection, longing; style: casual, conversational; average recording, quiet background; genuineness 4.5/6; vocal-burst blend 7.3/10; 11.6s, EN.
EN_5nsC7k4-UwE_W000011 · in -18.5 dBFS · gain -1.5 dB · emolia-00998
(fatigue exhaustion, impatience and irritability, anger · normal-paced, subdued, slightly relaxed, casual) Join the server, be a part of the RP and let's, let's branch off to, from patrols that are just starting around 7pm. We don't want to be doing that anymore. We want to have patrols start at 12 and end at 12 midnight or even go into the morning, maybe 3 o'clock in the morning (low mumble) at ESV.
full caption & clip details
A young adult masculine voice; delivery is subdued, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fatigue exhaustion, impatience and irritability, anger; style: casual, monologue; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 7.6/10; 18.1s, EN.
EN_5nsC7k4-UwE_W000012 · in -20.0 dBFS · gain -0.0 dB · emolia-00998
Emotional Numbness ↓  /  Painidentity −0.07 emotion 53 %   sc-AB2-k5 · #12

This chain comes from the two-sided rule: it only counts if both emotions move — Emotional Numbness down and Pain up — by at least 0.25 each.

The chain starts with Pain below average — 0.35, lower than 65 % of clips in this corpus — and ends with it clearly present at 0.72, higher than 72 % of clips in this corpus. That is a total rise of 0.37.

At the same time Emotional Numbness goes the other way, from 0.94 (higher than 94 % of clips in this corpus) to 0.62 (higher than 62 % of clips in this corpus), a change of -0.32. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are -0.24, then +0.24, then +0.20, then +0.17 — not a clean run: step 1 moves back the other way by 0.24 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.93 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.93 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.93), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 61 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.854 before conversion and 0.786 after — it fell by 0.068. Neighbour-to-neighbour the worst pair went 0.752 → 0.745. (The earlier render, with segment 1 left raw, scores 0.613 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Pain moved +0.369 in the original and +0.196 after conversion — 53 % of the delta retained. On the other named axis, Emotional Numbness, -0.325 became -0.212.

Quality. Mean predicted overall quality across the segments went 3.01 → 3.15 (+0.14) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.854 → 0.786 -0.068identity cos neighbours 0.752 → 0.745d_b rescored +0.369 → +0.196d_a rescored -0.325 → -0.212d_a mined -0.325d_b mined 0.369min_cos_consec (site) 0.9324min_cos_anchor (site) 0.9324dataset emolialang enspeaker EN_0wOGu4tXyyEtotal 59.2schain gain +1.6 dBseam step 2.5 dBcrossfades 150/100/150/100 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-bright, fairly smooth, balanced body, no background noise, normally alert, slightly relaxed, no disfluency, clear
(emotional numbness · normal-paced, fairly steady, moderate pitch range, newsreading) While their giving pattern matches that of other unions, public sector unions also concentrate contributions on members of Congress from both parties who sit on committees that deal with federal budgets and agencies
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: newsreading, formal; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.0/10; 11.1s, EN.
EN_0wOGu4tXyyE_W000016 · in -14.9 dBFS · gain -5.1 dB · emolia-00970
(sourness, emotional numbness · measured, steady, fairly narrow pitch, formal) Civil service Government agency Nationalization Political economy Privatization Public economics Public ownership Public sector business cases for projects Special purpose district State-owned enterprise
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, no audible breath; affect is neutral, neutral stance, slightly guarded; reads as sourness, emotional numbness; style: formal, authoritative; average recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.0/10; 20.6s, EN.
EN_0wOGu4tXyyE_W000018 · in -17.2 dBFS · gain -2.8 dB · emolia-00970
(emotional numbness · measured, steady, fairly narrow pitch, authoritative) Barlow, J. Rorick, J. K. and Wright, S. 2010. De facto privatization or a renewed role for the EU? Paying for Europe's healthcare infrastructure in a recession. Journal of the Royal Society of Medicine, 103-51-55.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, fairly narrow pitch, minimal breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: authoritative, formal; average recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.2/10; 16.1s, EN.
EN_0wOGu4tXyyE_W000022 · in -15.5 dBFS · gain -4.5 dB · emolia-00970
(normal-paced, fairly steady, moderate pitch range, formal) Lloyd G. Nigro, Decision Making in the Public Sector' 1984, Marcel Dekker Inc.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.7/10; 5.6s, EN.
EN_0wOGu4tXyyE_W000023 · in -14.5 dBFS · gain -5.5 dB · emolia-00970
(normal-paced, steady, moderate pitch range, formal) David G. Carnavale, Organizational Development in the Public Sector 2002, Westview PR
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.0/10; 6.6s, EN.
EN_0wOGu4tXyyE_W000024 · in -15.5 dBFS · gain -4.5 dB · emolia-00970
Concentration ↓  /  Embarrassmentidentity −0.01 emotion 32 %   sc-AB2-k5 · #13

This chain comes from the two-sided rule: it only counts if both emotions move — Concentration down and Embarrassment up — by at least 0.25 each.

The chain starts with Embarrassment around average — 0.58, higher than 58 % of clips in this corpus — and ends with it at the very top of the corpus at 0.92, higher than 92 % of clips in this corpus. That is a total rise of 0.34.

At the same time Concentration goes the other way, from 0.94 (higher than 94 % of clips in this corpus) to 0.55 (higher than 55 % of clips in this corpus), a change of -0.39. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.13, then +0.08, then -0.06, then +0.19 — not a clean run: step 3 moves back the other way by 0.06 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.95 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.95 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.95), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 74 s · bg · eurospeech

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.926 before conversion and 0.920 after — it fell by 0.005. Neighbour-to-neighbour the worst pair went 0.899 → 0.864. (The earlier render, with segment 1 left raw, scores 0.834 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Embarrassment moved +0.342 in the original and +0.108 after conversion — 32 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Concentration, -0.392 became -0.176.

Quality. Mean predicted overall quality across the segments went 3.01 → 3.20 (+0.19) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.926 → 0.920 -0.005identity cos neighbours 0.899 → 0.864d_b rescored +0.342 → +0.108d_a rescored -0.392 → -0.176d_a mined -0.394d_b mined 0.342min_cos_consec (site) 0.9451min_cos_anchor (site) 0.9510dataset eurospeechlang bgspeaker bulgaria_bulgaria_0_281120total 73.0schain gain +2.7 dBseam step 0.5 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, average recording, quiet background, normally alert, fairly steady, moderate pitch range
(concentration, sourness, contempt · normal-paced, slightly relaxed, almost no disfluency, monologue) Със Законопроекта се предвижда да се определя уникален идентификационен код на категоризираните места за настаняване, който ще се генерира от Националния туристически регистър и ще съдържа данни за категоризирания туристически обект.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, sourness, contempt; style: monologue, authoritative; average recording, quiet background; genuineness 0.8/6; vocal-burst blend 1.9/10; 15.2s, BG.
bulgaria_bulgaria_0_28112019_1373312_1388480 · in -32.8 dBFS · gain +12.8 dB · eurospeech-00071
(thankfulness gratitude · normal-paced, slightly relaxed, almost no disfluency, monologue) Създава се Национален регистър на туристическите забележителности, фестивали и събития, който да включва двата съществуващи регистъра – Регистъра на туристическите атракции и Регистъра на туристическите фестивали и събития.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as thankfulness gratitude; style: monologue, authoritative; average recording, quiet background; genuineness 0.8/6; vocal-burst blend 1.2/10; 13.3s, BG.
bulgaria_bulgaria_0_28112019_1388480_1401792 · in -33.5 dBFS · gain +13.5 dB · eurospeech-00071
(concentration, sourness, shame · brisk, neutral tension, some disfluency, monologue) тъй като систематичното място на правната норма е в Закона за туризма. Председателят на Комисията за защита на потребителите Димитър Маргаритов изрази подкрепа по Законопроекта и изказа мнение, че с предвидените разпоредби се оптимизира санкционният режим и се създават
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as concentration, sourness, shame; style: monologue, dramatic; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 5.4/10; 17.2s, BG.
bulgaria_bulgaria_0_28112019_1420096_1437312 · in -32.6 dBFS · gain +12.6 dB · eurospeech-00071
(shame, sourness, disgust · brisk, slightly relaxed, almost no disfluency, monologue) условия за по-добра координация между КЗП и Министерството на туризма. От сдружение „Българско ски училище“ настояват, по отношение на професионалната правоспособност и квалификацията на ски учителите, да не се препраща към наредбата по чл. 97, ал. 6 от Закона за физическото възпитание и спорта.
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as shame, sourness, disgust; style: monologue, authoritative; average recording, quiet background; genuineness 1.2/6; vocal-burst blend 3.4/10; 17.7s, BG.
bulgaria_bulgaria_0_28112019_1437312_1455024 · in -32.4 dBFS · gain +12.4 dB · eurospeech-00071
(embarrassment, relief, pride · normal-paced, slightly relaxed, some disfluency, monologue) От Българска туристическа камара изразиха притеснение, че с текстове от Законопроекта се ограничава участието им в експертните комисии.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as embarrassment, relief, pride; style: monologue, authoritative; average recording, quiet background; genuineness 1.3/6; vocal-burst blend 0.8/10; 10.4s, BG.
bulgaria_bulgaria_0_28112019_1455024_1465440 · in -33.6 dBFS · gain +13.6 dB · eurospeech-00071
Astonishment Surprise ↓  /  Intoxication Altered States of Consciousnessidentity +0.01 emotion 61 %   sc-AB2-k5 · #14

This chain comes from the two-sided rule: it only counts if both emotions move — Astonishment Surprise down and Intoxication Altered States of Consciousness up — by at least 0.25 each.

The chain starts with Intoxication Altered States of Consciousness around average — 0.57, higher than 57 % of clips in this corpus — and ends with it strongly present at 0.83, higher than 83 % of clips in this corpus. That is a total rise of 0.27.

At the same time Astonishment Surprise goes the other way, from 0.86 (higher than 86 % of clips in this corpus) to 0.52 (higher than 52 % of clips in this corpus), a change of -0.33. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.03, then +0.02, then +0.22, then -0.01 — not a clean run: step 4 moves back the other way by 0.01 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.81 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.81 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.81. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 25 s · zh · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.803 before conversion and 0.814 after — it rose by 0.012. Neighbour-to-neighbour the worst pair went 0.801 → 0.843. (The earlier render, with segment 1 left raw, scores 0.792 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Intoxication Altered States of Consciousness moved +0.266 in the original and +0.163 after conversion — 61 % of the delta retained. On the other named axis, Astonishment Surprise, -0.331 became -0.436.

Quality. Mean predicted overall quality across the segments went 2.81 → 2.86 (+0.05) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.803 → 0.814 +0.012identity cos neighbours 0.801 → 0.843d_b rescored +0.266 → +0.163d_a rescored -0.331 → -0.436d_a mined -0.331d_b mined 0.266min_cos_consec (site) 0.8092min_cos_anchor (site) 0.8092dataset emolialang zhspeaker ZH_B00081_S08363total 23.6schain gain +2.7 dBseam step 0.8 dBcrossfades 100/100/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, good recording, no background noise, normally alert, slightly relaxed, fairly steady
(brisk, formal, monologue) All the coaches of the train were packed into capacity ten minutes before it started.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 1.2/10; 5.1s, ZH.
ZH_B00081_S08363_W000036 · in -21.5 dBFS · gain +1.5 dB · emolia-04085
(normal-paced, authoritative, monologue) The most important thing in the olympics is not to win, but to participate.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: authoritative, monologue; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 2.6/10; 4.0s, ZH.
ZH_B00081_S08363_W000037 · in -21.3 dBFS · gain +1.3 dB · emolia-04085
(normal-paced, authoritative, monologue) The most important thing in the olympics is not to win, but to participate.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: authoritative, monologue; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 2.7/10; 4.0s, ZH.
ZH_B00081_S08363_W000038 · in -21.4 dBFS · gain +1.4 dB · emolia-04085
(brisk, authoritative, monologue) If the quality of your product meets with our customers approval, we will place regular orders.
full caption & clip details
An adult feminine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; no dominant emotion; style: authoritative, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 1.2/10; 5.7s, ZH.
ZH_B00081_S08363_W000039 · in -21.1 dBFS · gain +1.1 dB · emolia-04085
(brisk, authoritative, dramatic) If the quality of your product meets with our customers approval, we will place regular orders.
full caption & clip details
An adult feminine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, slightly bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; no dominant emotion; style: authoritative, dramatic; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 1.1/10; 5.7s, ZH.
ZH_B00081_S08363_W000040 · in -21.1 dBFS · gain +1.1 dB · emolia-04085
Contemplation ↓  /  Confusionidentity −0.01 emotion 46 %   sc-AB2-k5 · #15

This chain comes from the two-sided rule: it only counts if both emotions move — Contemplation down and Confusion up — by at least 0.25 each.

The chain starts with Confusion clearly present — 0.68, higher than 68 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.28.

At the same time Contemplation goes the other way, from 0.93 (higher than 93 % of clips in this corpus) to 0.62 (higher than 62 % of clips in this corpus), a change of -0.31. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.11, then +0.09, then -0.06, then +0.14 — not a clean run: step 3 moves back the other way by 0.06 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.91 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.89 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.91), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 47 s · ko · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.905 before conversion and 0.894 after — it fell by 0.011. Neighbour-to-neighbour the worst pair went 0.865 → 0.865. (The earlier render, with segment 1 left raw, scores 0.838 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Confusion moved +0.276 in the original and +0.128 after conversion — 46 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Contemplation, -0.309 became -0.162.

Quality. Mean predicted overall quality across the segments went 3.10 → 3.20 (+0.11) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.905 → 0.894 -0.011identity cos neighbours 0.865 → 0.865d_b rescored +0.276 → +0.128d_a rescored -0.309 → -0.162d_a mined -0.309d_b mined 0.276min_cos_consec (site) 0.8905min_cos_anchor (site) 0.9085dataset emolialang kospeaker KO_4QODsQ8zywEtotal 45.4schain gain +0.7 dBseam step 2.8 dBcrossfades 100/100/100/150 ms
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, average recording, quiet background, slightly relaxed, fairly steady, light breath
(contemplation, longing · measured, normally alert, frequent disfluency, didactic) 적어도 예배시간 30분 전에 도착해서 교제를 나누고 개인적으로 기도하고 찬송 연습에 참여한다.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation, longing; style: didactic, monologue; average recording, quiet background; genuineness 2.3/6; vocal-burst blend 0.9/10; 8.7s, KO.
KO_4QODsQ8zywE_W000162 · in -21.0 dBFS · gain +1.0 dB · emolia-03045
(measured, subdued, frequent disfluency, didactic) 이게 참 사실은 장로교 헌법 안에서 권장 사항이에요. (low mumble) 도착해서 기도로 준비하고, 또,
full caption & clip details
An adult masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: didactic, monologue; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 0.0/10; 11.1s, KO.
KO_4QODsQ8zywE_W000163 · in -20.5 dBFS · gain +0.5 dB · emolia-03045
(contemplation · measured, normally alert, some disfluency, conversational) 여러분, 준비 찬송은 없습니다. 준비, 찬송을 무슬 준비, 찬송 연습은 하는 거예요. 찬송 연습을 해서 같이 참여하고.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation; style: conversational, didactic; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 4.5/10; 12.3s, KO.
KO_4QODsQ8zywE_W000164 · in -18.1 dBFS · gain -1.9 dB · emolia-03045
(measured, normally alert, some disfluency, monologue) 너희들은 너리든 저쪽 가서 하고 우리는 이렇게. 그러지 않고 가족이 함께 앉아서 예배를 합니다.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; slurred, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, ASMR; average recording, quiet background; genuineness 2.6/6; vocal-burst blend 1.4/10; 6.2s, KO.
KO_4QODsQ8zywE_W000165 · in -21.1 dBFS · gain +1.1 dB · emolia-03045
(confusion, fear, distress · normal-paced, normally alert, some disfluency, didactic) 그리고 여기 이제 우리가 궁금한 거, 이곳저곳을 향하여 예배하거나 절하지 말고 바로 들어와 자리에 앉는다라는 설명이 뭐냐면요.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as confusion, fear, distress; style: didactic, authoritative; average recording, quiet background; genuineness 2.4/6; vocal-burst blend 5.8/10; 7.7s, KO.
KO_4QODsQ8zywE_W000166 · in -18.8 dBFS · gain -1.2 dB · emolia-03045
Concentration ↓  /  Impatience and Irritabilityidentity −0.09 emotion 117 %   sc-AB2-k5 · #16

This chain comes from the two-sided rule: it only counts if both emotions move — Concentration down and Impatience and Irritability up — by at least 0.25 each.

The chain starts with Impatience and Irritability clearly present — 0.71, higher than 71 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.27.

At the same time Concentration goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.71 (higher than 71 % of clips in this corpus), a change of -0.27. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.06, then -0.04, then +0.19, then +0.06 — not a clean run: step 2 moves back the other way by 0.04 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.88 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.89 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.88. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 69 s · zh · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.732 before conversion and 0.642 after — it fell by 0.090. Neighbour-to-neighbour the worst pair went 0.857 → 0.831. (The earlier render, with segment 1 left raw, scores 0.496 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Impatience and Irritability moved +0.270 in the original and +0.315 after conversion — 117 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Concentration, -0.266 became -0.198.

Quality. Mean predicted overall quality across the segments went 2.94 → 3.23 (+0.28) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.732 → 0.642 -0.090identity cos neighbours 0.857 → 0.831d_b rescored +0.270 → +0.315d_a rescored -0.266 → -0.198d_a mined -0.266d_b mined 0.270min_cos_consec (site) 0.8878min_cos_anchor (site) 0.8762dataset emolialang zhspeaker ZH_B00000_S07910total 68.1schain gain +2.7 dBseam step 1.1 dBcrossfades 100/100/100/100 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, brisk, normally alert, slightly relaxed, almost no disfluency, clear
(concentration, jealousy and envy, interest · fairly steady, monologue, narration) 但性格测试他就做到了。很多人看完这个性格分析之后,立马就找到了自己身边的人不给力的一个原因。原来你是什么什么性格的,所以你老是迟到。因为你是什么什么性格的,所以你注定是个内向的人。这个行业在高级一点就有点玄乎了,你捧着钱去排队都不一定轮得到,你就是那种专门给富人提供冥想的禅修班。
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, jealousy and envy, interest; style: monologue, narration; good recording, no background noise; genuineness 2.0/6; vocal-burst blend 7.0/10; 21.1s, ZH.
ZH_B00000_S07910_W000019 · in -18.7 dBFS · gain -1.3 dB · emolia-00019
(concentration, contempt, malevolence malice · fairly steady, monologue, cartoonish) 他们居然收费要十万以上。最神奇的是,这些客户去了之后,还觉得这个课程非常的神奇,因为他们难得有几天冷静下来思考人生,给他们人生里面发生的所有的事情都找到了原因。在这一行,你就靠短视频的流量入口,每天都会有大量的订单,会有大量的咨询服务。不要问我为什么,你去看看那些情感博主。
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as concentration, contempt, malevolence malice; style: monologue, cartoonish; good recording, quiet background; genuineness 2.1/6; vocal-burst blend 6.7/10; 21.7s, ZH.
ZH_B00000_S07910_W000020 · in -18.3 dBFS · gain -1.7 dB · emolia-00019
(fairly steady, monologue, authoritative) 现在的年轻人啊,宁愿养条狗也不结婚,生小孩花在宠物上的钱呢也是越来越多。十个年轻人中,可能就有六个正在养宠物。宠物市场正在疯狂的膨胀。
full caption & clip details
A young adult feminine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: monologue, authoritative; average recording, no background noise; genuineness 2.2/6; vocal-burst blend 5.2/10; 11.2s, ZH.
ZH_B00000_S07910_W000021 · in -18.3 dBFS · gain -1.7 dB · emolia-00019
(impatience and irritability, contempt · moderately variable, dramatic, authoritative) 未来五到十年会有更大的爆发期。未来宠物行业下面的一个细分赛道可能都能养活上百家企业了。
full caption & clip details
An adult feminine voice; delivery is normally alert, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as impatience and irritability, contempt; style: dramatic, authoritative; good recording, no background noise; genuineness 2.2/6; vocal-burst blend 5.2/10; 7.0s, ZH.
ZH_B00000_S07910_W000022 · in -17.8 dBFS · gain -2.2 dB · emolia-00019
(impatience and irritability, contempt, anger · moderately variable, dramatic, authoritative) 就是我这十八年的教学,确实帮到了一个又一个普通人逆袭成长,有一个最朴素的道理。就是我也不说我有多好,我就告诉你。
full caption & clip details
An adult feminine voice; delivery is normally alert, brisk, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as impatience and irritability, contempt, anger; style: dramatic, authoritative; average recording, no background noise; genuineness 2.5/6; vocal-burst blend 5.4/10; 7.8s, ZH.
ZH_B00000_S07910_W000023 · in -17.6 dBFS · gain -2.4 dB · emolia-00019
Relief ↓  /  Confusionidentity +0.03 emotion 96 %   sc-AB2-k5 · #17

This chain comes from the two-sided rule: it only counts if both emotions move — Relief down and Confusion up — by at least 0.25 each.

The chain starts with Confusion clearly present — 0.68, higher than 68 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.29.

At the same time Relief goes the other way, from 0.92 (higher than 92 % of clips in this corpus) to 0.58 (higher than 58 % of clips in this corpus), a change of -0.34. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.23, then -0.22, then +0.20, then +0.09 — not a clean run: step 2 moves back the other way by 0.22 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.86 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.84 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.86. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 72 s · zh · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.798 before conversion and 0.825 after — it rose by 0.027. Neighbour-to-neighbour the worst pair went 0.820 → 0.846. (The earlier render, with segment 1 left raw, scores 0.733 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Confusion moved +0.286 in the original and +0.274 after conversion — 96 % of the delta retained, which is essentially all of it. On the other named axis, Relief, -0.339 became -0.369.

Quality. Mean predicted overall quality across the segments went 3.19 → 3.36 (+0.17) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.798 → 0.825 +0.027identity cos neighbours 0.820 → 0.846d_b rescored +0.286 → +0.274d_a rescored -0.339 → -0.369d_a mined -0.339d_b mined 0.286min_cos_consec (site) 0.8444min_cos_anchor (site) 0.8646dataset emolialang zhspeaker ZH_B00004_S09712total 70.4schain gain +3.6 dBseam step 1.6 dBcrossfades 150/150/100/100 ms
Script — 5 chunks, 5 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, fairly smooth, balanced body, average recording, measured, fairly steady, moderate pitch range
(relief · normally alert, slightly relaxed, some disfluency, conversational) (low mumble) 秦始皇那个时候就有这个这么个法令啊,可能更早还是有的,到汉文帝这儿就给改了。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as relief; style: conversational, didactic; average recording, no background noise; genuineness 3.0/6; vocal-burst blend 4.4/10; 7.2s, ZH.
ZH_B00004_S09712_W000016 · in -18.3 dBFS · gain -1.7 dB · emolia-03314
(relief, confusion · normally alert, slightly relaxed, frequent disfluency, conversational) (ahem) (ahem) (low mumble) 啊,谁犯罪啊,谁惹事儿,你办谁就完了啊不要牵连,家里人,接着汉文帝又下了一道诏书。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as relief, confusion; style: conversational, monologue; average recording, quiet background; genuineness 4.1/6; vocal-burst blend 3.3/10; 9.4s, ZH.
ZH_B00004_S09712_W000017 · in -16.2 dBFS · gain -3.8 dB · emolia-03314
(normally alert, slightly relaxed, frequent disfluency, storytelling) (ahem) (low mumble) (ahem) 这个诏书太好了,就是开始啊救济各地的孤寡孤独啊。什么叫孤寡孤独啊?小朋友们。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; average clarity, frequent disfluency, moderate pitch range, audible breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: storytelling, monologue; average recording, no background noise; genuineness 2.6/6; vocal-burst blend 5.0/10; 10.0s, ZH.
ZH_B00004_S09712_W000018 · in -16.9 dBFS · gain -3.1 dB · emolia-03314
(intoxication altered states of consciousness, contemplation, concentration · subdued, relaxed, frequent disfluency, monologue) (low mumble) (low mumble) (low mumble) (wistful sigh) (low mumble) (low mumble) (ahem) 你知道吗?这个寡寡孤独其实是指四种人啊,这哪哪四种人呢?寡就是啊妻子去世的啊,这个这个这个那个男人啊,尤其是那个年老的人啊,寡是没了媳妇儿的寡呢呃就叫寡妇啊,是没了丈夫的人啊,孤儿独独没有儿的人,老年人叫毒。
full caption & clip details
An adult masculine voice; delivery is subdued, measured, relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as intoxication altered states of consciousness, contemplation, concentration; style: monologue, conversational; average recording, quiet background; genuineness 3.3/6; vocal-burst blend 5.2/10; 28.2s, ZH.
ZH_B00004_S09712_W000019 · in -19.6 dBFS · gain -0.5 dB · emolia-03314
(confusion, intoxication altered states of consciousness, astonishment surprise · normally alert, slightly relaxed, frequent disfluency, monologue) (low mumble) (low mumble) (low mumble) 高高兴兴的啊团团圆圆的一块儿工作一块儿啊,一块吃饭的啊,他呢就一个人日的苦。那么这些苦人呢啊汉文汉文帝特别为他们规定了。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as confusion, intoxication altered states of consciousness, astonishment surprise; style: monologue, casual; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 4.0/10; 16.2s, ZH.
ZH_B00004_S09712_W000020 · in -19.8 dBFS · gain -0.2 dB · emolia-03314
Infatuation ↓  /  Concentrationidentity −0.02 emotion 81 %   sc-AB2-k5 · #18

This chain comes from the two-sided rule: it only counts if both emotions move — Infatuation down and Concentration up — by at least 0.25 each.

The chain starts with Concentration around average — 0.50, right about the corpus median — and ends with it strongly present at 0.83, higher than 83 % of clips in this corpus. That is a total rise of 0.33.

At the same time Infatuation goes the other way, from 0.84 (higher than 84 % of clips in this corpus) to 0.51 (right about the corpus median), a change of -0.34. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.18, then +0.11, then -0.17, then +0.22 — not a clean run: step 3 moves back the other way by 0.17 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.86 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.89 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.86. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 42 s · zh · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.723 before conversion and 0.699 after — it fell by 0.024. Neighbour-to-neighbour the worst pair went 0.779 → 0.704. (The earlier render, with segment 1 left raw, scores 0.738 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.335 in the original and +0.271 after conversion — 81 % of the delta retained, which is most of it. On the other named axis, Infatuation, -0.336 became -0.203.

Quality. Mean predicted overall quality across the segments went 3.02 → 3.11 (+0.09) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.723 → 0.699 -0.024identity cos neighbours 0.779 → 0.704d_b rescored +0.335 → +0.271d_a rescored -0.336 → -0.203d_a mined -0.336d_b mined 0.335min_cos_consec (site) 0.8900min_cos_anchor (site) 0.8613dataset emolialang zhspeaker ZH_B00007_S06562total 41.0schain gain +1.0 dBseam step 1.0 dBcrossfades 150/150/100/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, no background noise, normally alert, slightly relaxed, fairly steady
(normal-paced, formal, authoritative) 尽管很明显,格林伯格掌管着集团内部所有的事情。
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 2.7/10; 3.5s, ZH.
ZH_B00007_S06562_W000102 · in -18.5 dBFS · gain -1.5 dB · emolia-03344
(confusion · measured, narration, monologue) 而且格林伯格和高盛的顶层投资银行家们有密切的关系。但是一位钱高盛银行家回忆说,每次开始筹备新的交易时。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as confusion; style: narration, monologue; average recording, no background noise; genuineness 0.6/6; vocal-burst blend 5.3/10; 9.4s, ZH.
ZH_B00007_S06562_W000103 · in -19.5 dBFS · gain -0.5 dB · emolia-03344
(measured, monologue, formal) 参与工作的年轻银行家们会发现,集团在内部控制上到处存在缺陷。在尽职调查的电话中,他们总是对我们的提问采取迁回战术。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 3.6/10; 10.3s, ZH.
ZH_B00007_S06562_W000104 · in -19.4 dBFS · gain -0.7 dB · emolia-03344
(normal-paced, authoritative, formal) 那我们只能去找高盛中和美国国际集团具有深层关系的人。
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: authoritative, formal; good recording, no background noise; genuineness 1.3/6; vocal-burst blend 4.8/10; 4.0s, ZH.
ZH_B00007_S06562_W000105 · in -18.1 dBFS · gain -1.9 dB · emolia-03344
(measured, monologue, formal) 自身管理者都不愿意听到任何关于美国国际集团内部问题的消息,交易基本都能顺利进行下去。接下来的问题是,美国国际集团承担的风险,尽管公司的规模已经很大。
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, formal; average recording, no background noise; genuineness 0.9/6; vocal-burst blend 4.3/10; 14.5s, ZH.
ZH_B00007_S06562_W000106 · in -19.3 dBFS · gain -0.7 dB · emolia-03344
Pain ↓  /  Concentrationidentity −0.09 emotion 85 %   sc-AB2-k5 · #19

This chain comes from the two-sided rule: it only counts if both emotions move — Pain down and Concentration up — by at least 0.25 each.

The chain starts with Concentration around average — 0.57, higher than 57 % of clips in this corpus — and ends with it strongly present at 0.89, higher than 89 % of clips in this corpus. That is a total rise of 0.32.

At the same time Pain goes the other way, from 0.85 (higher than 85 % of clips in this corpus) to 0.20 (lower than 80 % of clips in this corpus), a change of -0.65. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.02, then +0.10, then +0.13, then +0.06 — a plateau around step 1, where it barely moves.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.86 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.84 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.86. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 35 s · zh · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.800 before conversion and 0.709 after — it fell by 0.091. Neighbour-to-neighbour the worst pair went 0.815 → 0.702. (The earlier render, with segment 1 left raw, scores 0.690 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.320 in the original and +0.272 after conversion — 85 % of the delta retained, which is most of it. On the other named axis, Pain, -0.653 became -0.749.

Quality. Mean predicted overall quality across the segments went 2.88 → 3.04 (+0.16) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.800 → 0.709 -0.091identity cos neighbours 0.815 → 0.702d_b rescored +0.320 → +0.272d_a rescored -0.653 → -0.749d_a mined -0.653d_b mined 0.320min_cos_consec (site) 0.8418min_cos_anchor (site) 0.8559dataset emolialang zhspeaker ZH_B00081_S02600total 34.0schain gain +0.2 dBseam step 2.9 dBcrossfades 100/100/150/100 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, good recording, no background noise, normally alert, slightly relaxed, fairly steady
(normal-paced, no disfluency, storytelling, formal) 看来,俄国人所拥有的飞机和坦克的数量很大。
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, neutral openness; no dominant emotion; style: storytelling, formal; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 3.6/10; 3.9s, ZH.
ZH_B00081_S02600_W000063 · in -18.2 dBFS · gain -1.8 dB · emolia-04082
(fast, no disfluency, dramatic, storytelling) 到目前为止,仅击毁的飞机就有三千五百架运油车一千辆。
full caption & clip details
A young adult feminine voice; delivery is normally alert, fast, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: dramatic, storytelling; good recording, no background noise; genuineness 1.3/6; vocal-burst blend 3.9/10; 5.8s, ZH.
ZH_B00081_S02600_W000064 · in -17.4 dBFS · gain -2.6 dB · emolia-04082
(measured, no disfluency, storytelling, narration) 如果俄国人组织领导的好,这场战斗将是难分胜负的。
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: storytelling, narration; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 4.1/10; 4.4s, ZH.
ZH_B00081_S02600_W000065 · in -17.6 dBFS · gain -2.4 dB · emolia-04082
(measured, no disfluency, monologue, narration) 纵观迄今为止的情况,可以断言人们是在与一群野兽打仗。你知道为什么我们才抓了那么一点俘虏?
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, narration; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 1.5/10; 8.7s, ZH.
ZH_B00081_S02600_W000066 · in -17.4 dBFS · gain -2.6 dB · emolia-04082
(normal-paced, almost no disfluency, monologue, narration) 这是因为俄国人受到了他们的政治委员们的煽动,他们听了捏造的有关我们不人道的惨文谎说。如果他们被我们抓住,他们就会受到这种不人道的待遇。
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: monologue, narration; good recording, no background noise; genuineness 0.9/6; vocal-burst blend 2.7/10; 11.9s, ZH.
ZH_B00081_S02600_W000067 · in -17.4 dBFS · gain -2.6 dB · emolia-04082
Emotional Numbness ↓  /  Concentrationidentity −0.08 emotion 101 %   sc-AB2-k5 · #20

This chain comes from the two-sided rule: it only counts if both emotions move — Emotional Numbness down and Concentration up — by at least 0.25 each.

The chain starts with Concentration below average — 0.41, lower than 59 % of clips in this corpus — and ends with it strongly present at 0.75, higher than 75 % of clips in this corpus. That is a total rise of 0.35.

At the same time Emotional Numbness goes the other way, from 0.88 (higher than 88 % of clips in this corpus) to 0.57 (higher than 57 % of clips in this corpus), a change of -0.31. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.11, then +0.23, then -0.09, then +0.09 — not a clean run: step 3 moves back the other way by 0.09 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.90 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.90 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.90. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 48 s · en · emolia

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.785 before conversion and 0.703 after — it fell by 0.082. Neighbour-to-neighbour the worst pair went 0.859 → 0.772. (The earlier render, with segment 1 left raw, scores 0.512 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.347 in the original and +0.349 after conversion — 101 % of the delta retained, which is essentially all of it. On the other named axis, Emotional Numbness, -0.312 became -0.324.

Quality. Mean predicted overall quality across the segments went 2.90 → 3.04 (+0.14) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.785 → 0.703 -0.082identity cos neighbours 0.859 → 0.772d_b rescored +0.347 → +0.349d_a rescored -0.312 → -0.324d_a mined -0.312d_b mined 0.346min_cos_consec (site) 0.8954min_cos_anchor (site) 0.8954dataset emolialang enspeaker EN_E_Hezvw0RSytotal 46.9schain gain +1.4 dBseam step 0.9 dBcrossfades 150/100/150/100 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(formal, authoritative) Tindal took the issue to the High Court, where his expulsion was declared illegal
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: formal, authoritative; good recording, no background noise; genuineness 0.3/6; vocal-burst blend 1.0/10; 4.2s, EN.
EN_E_Hezvw0RSy_W000082 · in -14.3 dBFS · gain -5.7 dB · emolia-00691
(emotional numbness · formal, newsreading) In frustration at their inability to eject Tyndall and the Tyndallites, Reid and his supporters split from the NF to form the National Party
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 9.7s, EN.
EN_E_Hezvw0RSy_W000083 · in -14.3 dBFS · gain -5.7 dB · emolia-00691
(emotional numbness, pain · formal, newsreading) By the end of February 1976, 29 NF branches and groups had defected to the NP, although 101 remained loyal.In February 1976, Tyndall was restored as the NF leader
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, pain; style: formal, newsreading; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 12.1s, EN.
EN_E_Hezvw0RSy_W000084 · in -14.4 dBFS · gain -5.6 dB · emolia-00691
(malevolence malice, anger, bitterness · formal, newsreading) The party then capitalised on public anger at the government's agreement to accept Malawian Asian refugees, holding demonstrations to protest the migrants' arrival in the UK
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as malevolence malice, anger, bitterness; style: formal, newsreading; good recording, no background noise; genuineness 0.2/6; vocal-burst blend 0.0/10; 9.6s, EN.
EN_E_Hezvw0RSy_W000085 · in -14.5 dBFS · gain -5.5 dB · emolia-00691
(newsreading, formal) After a resurgence in fortunes for the party in London at the 1977 GLC election, where it improved on its October 1974 general election result, it planned further marches in the city
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; no dominant emotion; style: newsreading, formal; good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.0/10; 12.0s, EN.
EN_E_Hezvw0RSy_W000086 · in -14.3 dBFS · gain -5.7 dB · emolia-00691