c-eurospeech-AB2 — voice-corrected

Corpus eurospeech in isolation, rule AB2.

This is the voice-conversion-corrected version of this tier. The original, uncorrected page is still there and unchanged: t_c-eurospeech-AB2.html. Every card below carries both renders so you can switch between them without leaving the page. What was done, and what it measured →
20chains converted
61segments re-voiced
0.449 → 0.834median worst-to-anchor identity cosine
92 %median emotion delta retained
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken.

(underlined, plain · delivery, style) — the tag before the words. Emotions first, then how it is delivered. Underlined descriptors are the ones that change across this chain — anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words are a real non-speech sound, printed where it happens.

The full generated caption for any clip is under “full caption & clip details”. Its perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.

The tags and captions describe the ORIGINAL clips — they are the corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the converted audio are the ones printed in each card's own paragraph and chips.
Jealousy and Envy ↓  /  Contemplationidentity +0.37 emotion 81 %   c-eurospeech-AB2 · #1

This chain comes from the two-sided rule: it only counts if both emotions move — Jealousy and Envy down and Contemplation up — by at least 0.25 each.

The chain starts with Contemplation around average — 0.57, higher than 57 % of clips in this corpus — and ends with it at the very top of the corpus at 0.95, higher than 95 % of clips in this corpus. That is a total rise of 0.37.

At the same time Jealousy and Envy goes the other way, from 1.00 (virtually no clip in this corpus scores higher) to 0.60 (higher than 60 % of clips in this corpus), a change of -0.40. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.17, then +0.19, then +0.01 — a plateau around step 3, where it barely moves.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.15 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.24 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.15, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 56 s · en · eurospeech

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.201 before conversion and 0.573 after — it rose by 0.373. Neighbour-to-neighbour the worst pair went 0.256 → 0.573. (The earlier render, with segment 1 left raw, scores 0.607 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contemplation moved +0.373 in the original and +0.302 after conversion — 81 % of the delta retained, which is most of it. On the other named axis, Jealousy and Envy, -0.397 became -0.309.

Quality. Mean predicted overall quality across the segments went 2.92 → 3.27 (+0.35) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.201 → 0.573 +0.373identity cos neighbours 0.256 → 0.573d_b rescored +0.373 → +0.302d_a rescored -0.397 → -0.309d_a mined -0.397d_b mined 0.374min_cos_consec (site) 0.2375min_cos_anchor (site) 0.1543dataset eurospeechlang enspeaker uk_uk_0_29112024total 54.5schain gain +3.6 dBseam step 0.6 dBcrossfades 150/150/150 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a young adult feminine voice · average recording
(jealousy and envy, contempt, sourness · brisk, energised, neutral tension, dramatic) are talking about this in the television and radio studios, they should think of those in my constituency who have poor English, or the woman who came to see me a month ago with terrible pain in her gall bladder.
full caption & clip details
A young adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; clear, almost no disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, fairly guarded; reads as jealousy and envy, contempt, sourness; style: dramatic, ranting; average recording, quiet background; genuineness 1.4/6; vocal-burst blend 3.9/10; 11.8s, EN.
uk_uk_0_29112024_8063824_8075632 · in -24.4 dBFS · gain +4.4 dB · eurospeech-00840
(embarrassment, emotional numbness, doubt · measured, subdued, neutral tension, narration) (low mumble) by a personal friend and constituent: “I apologise for adding to the thousands of emails you will be receiving. I just wanted to tell you why I oppose the right to die Bill. I know you are aware of the experience I had when my husband was dying.
full caption & clip details
An elderly masculine voice; delivery is subdued, measured, neutral tension, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, little disfluency, fairly narrow pitch, light breath; affect is mildly negative, neutral stance, slightly guarded; reads as embarrassment, emotional numbness, doubt; style: narration, monologue; average recording, quiet background; genuineness 1.6/6; vocal-burst blend 2.3/10; 16.8s, EN.
uk_uk_0_29112024_8116544_8133344 · in -28.8 dBFS · gain +8.8 dB · eurospeech-00840
(bitterness, distress, sadness · slow, very low-energy, slightly relaxed, narration) In hospital we had a dreadful experience because they had no end-of-life care and he suffered. Once in the Hospice it was a different story and he received the loving care he rightly deserved.
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, slow, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; clear, almost no disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, neutral openness; reads as bitterness, distress, sadness; style: narration, storytelling; average recording, no background noise; genuineness 0.7/6; vocal-burst blend 0.0/10; 13.6s, EN.
uk_uk_0_29112024_8133344_8146896 · in -30.5 dBFS · gain +10.5 dB · eurospeech-00840
(contemplation, fatigue exhaustion, disappointment · slow, very low-energy, slightly relaxed, narration) My argument is that, instead of assisted dying, we should be spending much more money on end-of-life care and funding the wonderful Hospice movement. Thank you for reading this.”
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, slow, slightly relaxed, steady; timbre is slightly cool, slightly dark, rough, very full; clear, little disfluency, fairly narrow pitch, normal breath; affect is negative, slightly dominant, fairly guarded; reads as contemplation, fatigue exhaustion, disappointment; style: narration, monologue; average recording, quiet background; genuineness 0.4/6; vocal-burst blend 0.0/10; 12.9s, EN.
uk_uk_0_29112024_8146896_8159824 · in -30.5 dBFS · gain +10.5 dB · eurospeech-00840
Relief ↓  /  Angeridentity −0.03 emotion 67 %   c-eurospeech-AB2 · #2

This chain comes from the two-sided rule: it only counts if both emotions move — Relief down and Anger up — by at least 0.25 each.

The chain starts with Anger clearly present — 0.67, higher than 67 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.30.

At the same time Relief goes the other way, from 0.99 (higher than 99 % of clips in this corpus) to 0.72 (higher than 72 % of clips in this corpus), a change of -0.27. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.21, then +0.09 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.93 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.93 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.93), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 49 s · lt · eurospeech

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.889 before conversion and 0.861 after — it fell by 0.028. Neighbour-to-neighbour the worst pair went 0.926 → 0.881. (The earlier render, with segment 1 left raw, scores 0.649 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Anger moved +0.296 in the original and +0.197 after conversion — 67 % of the delta retained. On the other named axis, Relief, -0.268 became -0.233.

Quality. Mean predicted overall quality across the segments went 3.28 → 3.41 (+0.13) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.889 → 0.861 -0.028identity cos neighbours 0.926 → 0.881d_b rescored +0.296 → +0.197d_a rescored -0.268 → -0.233d_a mined -0.268d_b mined 0.296min_cos_consec (site) 0.9312min_cos_anchor (site) 0.9302dataset eurospeechlang ltspeaker lithuania_lithuania_10_221total 48.0schain gain +0.7 dBseam step 1.2 dBcrossfades 150/150 ms
Script — 3 chunks, 2 with a non-speech sound
Unchanged across all 3 clips: a middle-aged masculine voice · neutral-toned, neutral-bright, balanced body, average recording, slightly relaxed, fairly steady, somewhat unclear, normal breath
(relief, concentration, thankfulness gratitude · measured, normally alert, frequent disfluency, monologue) nes pradinis (low mumble) pasiūlymas vėlgi buvo toks, ką aš vadinu kavaleristine ataka, kai buvo siūloma didinti tik lengvų gėrimų, t. y. alaus ir vyno, akcizą per 100 % 2017 metais, paliekant stiprių gėrimų faktiškai nedidinamus.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as relief, concentration, thankfulness gratitude; style: monologue; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 1.4/10; 19.1s, LT.
lithuania_lithuania_10_22122016_5509568_5528639 · in -22.4 dBFS · gain +2.4 dB · eurospeech-01833
(pride, triumph, relief · normal-paced, normally alert, frequent disfluency, authoritative) tačiau klausimas, į kurį aš neturiu atsakymo: ar iš tikrųjų (low mumble) tokiam didinimui yra tinkamai pasirengusi Vyriausybė ir kitos institucijos, nes (low mumble) pats svarstymas demonstravo tai, kad (low mumble)
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as pride, triumph, relief; style: authoritative; average recording, some background noise; genuineness 4.0/6; vocal-burst blend 2.4/10; 16.8s, LT.
lithuania_lithuania_10_22122016_5554255_5571008 · in -19.6 dBFS · gain -0.4 dB · eurospeech-01833
(anger, disgust, malevolence malice · normal-paced, energised, almost no disfluency, authoritative) labai atidžiai stebėti situaciją, nes tikrai ir įvairios suinteresuotos grupės, kad ši programa būtų diskredituota, stengsis pasinaudoti bet kokiomis klaidomis.
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, almost no disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, fairly guarded; reads as anger, disgust, malevolence malice; style: authoritative, monologue; average recording, quiet background; genuineness 1.1/6; vocal-burst blend 2.2/10; 12.6s, LT.
lithuania_lithuania_10_22122016_5587360_5599936 · in -17.6 dBFS · gain -2.4 dB · eurospeech-01833
Anger ↓  /  Thankfulness Gratitudeidentity +0.42 emotion 8 %   c-eurospeech-AB2 · #3

This chain comes from the two-sided rule: it only counts if both emotions move — Anger down and Thankfulness Gratitude up — by at least 0.25 each.

The chain starts with Thankfulness Gratitude around average — 0.55, higher than 55 % of clips in this corpus — and ends with it at the very top of the corpus at 0.91, higher than 91 % of clips in this corpus. That is a total rise of 0.37.

At the same time Anger goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.52 (higher than 52 % of clips in this corpus), a change of -0.45. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.24, then +0.07, then +0.07 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.41 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.28 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.41, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 64 s · no · eurospeech

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.391 before conversion and 0.807 after — it rose by 0.416. Neighbour-to-neighbour the worst pair went 0.286 → 0.763. (The earlier render, with segment 1 left raw, scores 0.769 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Thankfulness Gratitude moved +0.369 in the original and +0.031 after conversion — 8 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Anger, -0.455 became -0.354.

Quality. Mean predicted overall quality across the segments went 3.10 → 3.44 (+0.34) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.391 → 0.807 +0.416identity cos neighbours 0.286 → 0.763d_b rescored +0.369 → +0.031d_a rescored -0.455 → -0.354d_a mined -0.455d_b mined 0.369min_cos_consec (site) 0.2783min_cos_anchor (site) 0.4055dataset eurospeechlang nospeaker norway_10591-1total 62.7schain gain +2.3 dBseam step 3.7 dBcrossfades 150/150/150 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a middle-aged masculine voice · neutral-bright, balanced body, quiet background, average clarity, wide pitch range
(anger, awe, interest · measured, energised, neutral tension, cartoonish) det bedre. Når Kari Henriksen og jeg viser til den manglende tilliten, er det ikke vi som uttrykker det, det er folk som uttrykker det. Den manglende rettssikkerheten er det ikke vi som uttrykker, det er folk som uttrykker det. Vi satt selv i høring her
full caption & clip details
A middle-aged masculine voice; delivery is energised, measured, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, frequent disfluency, wide pitch range, audible breath; affect is neutral, slightly dominant, slightly guarded; reads as anger, awe, interest; style: cartoonish, monologue; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 1.7/10; 18.1s, NO.
norway_10591-1_10844288_10862416 · in -25.5 dBFS · gain +5.5 dB · eurospeech-02230
(triumph, pride · normal-paced, normally alert, slightly relaxed, authoritative) å gjøre folks liv bedre, og da hjelper engasjement, men det hjel­ per mest med handling. Kristin Ørmen Johnsen
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as triumph, pride; style: authoritative, dramatic; good recording, quiet background; genuineness 2.3/6; vocal-burst blend 0.2/10; 15.0s, NO.
norway_10591-1_10879008_10894008 · in -29.4 dBFS · gain +9.4 dB · eurospeech-02230
(longing, contemplation · measured, very low-energy, slightly relaxed, playful) Ja, det er jo handling (ahem) vi gjør nå. Vi vedtar (wistful sigh) fem (childlike giggle) viktige punkter i en barnevernslov – det er handling – og det er punkter som vi i komiteen faktisk er enige om.
full caption & clip details
A middle-aged somewhat feminine voice; delivery is very low-energy, measured, slightly relaxed, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, balanced body; average clarity, frequent disfluency, wide pitch range, light breath; affect is positive, neutral stance, neutral openness; reads as longing, contemplation; style: playful, monologue; average recording, quiet background; genuineness 2.8/6; vocal-burst blend 1.2/10; 14.8s, NO.
norway_10591-1_10904244_10919088 · in -24.9 dBFS · gain +4.9 dB · eurospeech-02230
(thankfulness gratitude · normal-paced, normally alert, slightly relaxed, monologue) derfor skal denne loven ha virkning fra 1. januar 2021. Men statsråden har også sagt her, og komiteen har bedt om, at man vurderer tiltak allerede nå for dem som vil falle inn under ny lov.
full caption & clip details
A middle-aged feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as thankfulness gratitude; style: monologue, didactic; average recording, quiet background; genuineness 2.6/6; vocal-burst blend 0.7/10; 15.3s, NO.
norway_10591-1_10930896_10946176 · in -27.5 dBFS · gain +7.5 dB · eurospeech-02230
Doubt ↓  /  Angeridentity +0.80 emotion 50 %   c-eurospeech-AB2 · #4

This chain comes from the two-sided rule: it only counts if both emotions move — Doubt down and Anger up — by at least 0.25 each.

The chain starts with Anger clearly present — 0.65, higher than 65 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.32.

At the same time Doubt goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.49 (lower than 51 % of clips in this corpus), a change of -0.48. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.21, then +0.02, then +0.09 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.04 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.16 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.04, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 47 s · en · eurospeech

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.019 before conversion and 0.816 after — it rose by 0.797. Neighbour-to-neighbour the worst pair went 0.142 → 0.796. (The earlier render, with segment 1 left raw, scores 0.665 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Anger moved +0.316 in the original and +0.158 after conversion — 50 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Doubt, -0.476 became -0.490.

Quality. Mean predicted overall quality across the segments went 2.99 → 3.17 (+0.18) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.019 → 0.816 +0.797identity cos neighbours 0.142 → 0.796d_b rescored +0.316 → +0.158d_a rescored -0.476 → -0.490d_a mined -0.476d_b mined 0.316min_cos_consec (site) 0.1555min_cos_anchor (site) 0.0387dataset eurospeechlang enspeaker uk_uk_19_11032016total 46.0schain gain +4.1 dBseam step 1.5 dBcrossfades 100/150/150 ms
Script — 4 chunks, 3 with a non-speech sound
Unchanged across all 4 clips: a middle-aged somewhat feminine voice · neutral-bright, fairly smooth, moderate pitch range
(doubt, disappointment, sourness · normal-paced, normally alert, neutral tension, casual) is talking about the international development budget and prisons abroad. In the very uncertain world in which we now
full caption & clip details
A middle-aged somewhat feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, frequent disfluency, moderate pitch range, normal breath; affect is mildly positive, neutral stance, neutral openness; reads as doubt, disappointment, sourness; style: casual, conversational; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 0.7/10; 10.3s, EN.
uk_uk_19_11032016_3991632_4001920 · in -26.5 dBFS · gain +6.5 dB · eurospeech-00915
(fear, distress, disappointment · normal-paced, normally alert, slightly relaxed, whispered) does he not agree that it is good that (ahem) our Government are spending money on strengthening the legal systems in these countries so that they can deal with their own prisoners?
full caption & clip details
A middle-aged somewhat feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as fear, distress, disappointment; style: whispered, formal; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 0.2/10; 11.3s, EN.
uk_uk_19_11032016_4001920_4013248 · in -26.9 dBFS · gain +6.9 dB · eurospeech-00915
(brisk, normally alert, slightly relaxed, dramatic) fact refers to “any foreign national convicted in any court of law”. I fear that my hon. Friend the Member for Christchurch (Mr Chope) may need to introduce a new Bill (ahem) if we are to seek savings in translation services,
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; no dominant emotion; style: dramatic, monologue; good recording, quiet background; genuineness 1.3/6; vocal-burst blend 3.1/10; 11.3s, EN.
uk_uk_19_11032016_4031968_4043312 · in -24.4 dBFS · gain +4.4 dB · eurospeech-00915
(anger, bitterness, sourness · brisk, energised, neutral tension, conversational) 20 or 30 years ago, when they were under the Soviet yoke, they were not able to travel to this country to work? As I have said, they and nationals from all the other countries that have been mentioned do a tremendous amount of good (low mumble) for the British economy.
full caption & clip details
An adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as anger, bitterness, sourness; style: conversational, dramatic; good recording, quiet background; genuineness 2.9/6; vocal-burst blend 5.5/10; 13.6s, EN.
uk_uk_19_11032016_4099808_4113424 · in -23.2 dBFS · gain +3.2 dB · eurospeech-00915
Concentration ↓  /  Reliefidentity +0.48 emotion 186 %   c-eurospeech-AB2 · #5

This chain comes from the two-sided rule: it only counts if both emotions move — Concentration down and Relief up — by at least 0.25 each.

The chain starts with Relief clearly present — 0.63, higher than 63 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.36.

At the same time Concentration goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.67 (higher than 67 % of clips in this corpus), a change of -0.31. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.14, then +0.19, then -0.11, then +0.14 — not a clean run: step 3 moves back the other way by 0.11 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.21 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.23 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.21, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 74 s · lt · eurospeech

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.218 before conversion and 0.702 after — it rose by 0.484. Neighbour-to-neighbour the worst pair went 0.301 → 0.702. (The earlier render, with segment 1 left raw, scores 0.385 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Relief moved +0.358 in the original and +0.666 after conversion — 186 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Concentration, -0.306 became -0.222.

Quality. Mean predicted overall quality across the segments went 2.82 → 3.33 (+0.51) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.218 → 0.702 +0.484identity cos neighbours 0.301 → 0.702d_b rescored +0.358 → +0.666d_a rescored -0.306 → -0.222d_a mined -0.305d_b mined 0.358min_cos_consec (site) 0.2324min_cos_anchor (site) 0.2135dataset eurospeechlang ltspeaker lithuania_lithuania_15_160total 72.7schain gain +1.1 dBseam step 0.5 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: a middle-aged masculine voice · quiet background, normally alert, somewhat unclear
(concentration, anger, pride · normal-paced, neutral tension, moderately variable, cartoonish) Ar jūs nemanote, kad jūsų tarnyba turėtų ketvirčiais ar pusmečiais per regioninę spaudą jūsų vykdomą monitoringą apie šituos aprašus skelbti visuomenei, ar tai pasitvirtino šie straipsniai, ar ne? Jūsų nuomonė?
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is slightly cool, slightly dark, slightly rough, thin; somewhat unclear, some disfluency, wide pitch range, audible breath; affect is neutral, slightly dominant, fairly guarded; reads as concentration, anger, pride; style: cartoonish, authoritative; below-average recording, quiet background; genuineness 3.1/6; vocal-burst blend 5.4/10; 17.9s, LT.
lithuania_lithuania_15_16062009_6947904_6965808 · in -11.4 dBFS · gain -8.6 dB · eurospeech-01877
(malevolence malice, triumph, fear · normal-paced, slightly relaxed, steady, monologue) bet kokia informacija, tarp jų ir spaudoje pasirodanti informacija, yra labai svarbi Specialiųjų tyrimų tarnybai vykdant savo funkcijas.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as malevolence malice, triumph, fear; style: monologue, authoritative; average recording, quiet background; genuineness 1.1/6; vocal-burst blend 0.0/10; 12.3s, LT.
lithuania_lithuania_15_16062009_6978960_6991233 · in -21.8 dBFS · gain +1.8 dB · eurospeech-01877
(thankfulness gratitude, pride, triumph · normal-paced, slightly relaxed, steady, monologue) Nacionalinio saugumo ir gynybos komitetas šių metų balandžio 22 d. svarstė Specialiųjų tyrimų tarnybos 2008 metų veiklos ataskaitą
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as thankfulness gratitude, pride, triumph; style: monologue; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 0.4/10; 10.4s, LT.
lithuania_lithuania_15_16062009_7032480_7042880 · in -16.9 dBFS · gain -3.1 dB · eurospeech-01877
(triumph, pride, bitterness · normal-paced, slightly relaxed, fairly steady, monologue) ir pritarė jai bendru sutarimu bei pateikė Seimo nutarimo ‘Dėl Specialiųjų tyrimų tarnybos 2008 metų veiklos ataskaitos’ projektą. Nutarimo projekto 1 straipsniu siūloma Seimui pritarti tarnybos 2008 metų veiklos ataskaitai, o 2 straipsniu pabrėžti
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, normal breath; affect is neutral, slightly dominant, fairly guarded; reads as triumph, pride, bitterness; style: monologue, authoritative; average recording, quiet background; genuineness 1.9/6; vocal-burst blend 2.1/10; 19.4s, LT.
lithuania_lithuania_15_16062009_7042880_7062304 · in -15.7 dBFS · gain -4.3 dB · eurospeech-01877
(relief, sourness, triumph · measured, slightly relaxed, fairly steady, monologue) (ahem) (low mumble) tarnybai konkretūs pasiūlymai, kurie stiprintų kovą su korupcija, darytų tarnybos veiklą efektyvesnę. Aš visų siūlymų neminėsiu, jūs juos turite.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as relief, sourness, triumph; style: monologue; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 2.2/10; 13.5s, LT.
lithuania_lithuania_15_16062009_7081983_7095520 · in -14.3 dBFS · gain -5.7 dB · eurospeech-01877
Shame ↓  /  Emotional Numbnessidentity −0.02 emotion 138 %   c-eurospeech-AB2 · #6

This chain comes from the two-sided rule: it only counts if both emotions move — Shame down and Emotional Numbness up — by at least 0.25 each.

The chain starts with Emotional Numbness clearly present — 0.68, higher than 68 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.31.

At the same time Shame goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.70 (higher than 70 % of clips in this corpus), a change of -0.27. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.06, then +0.24 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.94 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.91 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.94), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 40 s · el · eurospeech

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.869 before conversion and 0.852 after — it fell by 0.017. Neighbour-to-neighbour the worst pair went 0.896 → 0.835. (The earlier render, with segment 1 left raw, scores 0.625 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.305 in the original and +0.420 after conversion — 138 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Shame, -0.272 became -0.054.

Quality. Mean predicted overall quality across the segments went 3.05 → 3.34 (+0.28) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.869 → 0.852 -0.017identity cos neighbours 0.896 → 0.835d_b rescored +0.305 → +0.420d_a rescored -0.272 → -0.054d_a mined -0.272d_b mined 0.305min_cos_consec (site) 0.9120min_cos_anchor (site) 0.9361dataset eurospeechlang elspeaker greece_olomeleia-20210713atotal 39.6schain gain +2.5 dBseam step 0.2 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a middle-aged masculine voice · neutral-toned, balanced body, average recording, quiet background, measured, fairly steady, somewhat unclear, normal breath
(shame, triumph, thankfulness gratitude · subdued, neutral tension, frequent disfluency, monologue) μέσα στο έτος, όπως και η νέα περιφερειακή της Θεσσαλονίκης, το γνωστό «Flyover», και δύο οδικές συνδέσεις για τις οποίες είχα δεσμευτεί προσωπικά, ο δρόμος Θεσσαλονίκη - Έδεσσα και ο δρόμος Δράμα - Αμφίπολη,
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, neutral tension, fairly steady; timbre is neutral-toned, slightly dark, rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, slightly dominant, fairly guarded; reads as shame, triumph, thankfulness gratitude; style: monologue, authoritative; average recording, quiet background; genuineness 2.2/6; vocal-burst blend 1.0/10; 18.3s, EL.
greece_olomeleia-20210713a_9309456_9327776 · in -23.8 dBFS · gain +3.8 dB · eurospeech-00790
(contempt · normally alert, neutral tension, some disfluency, authoritative) τρεις αρτηρίες με συνολικό κόστος κοντά στο 1 δισεκατομμύριο, που απαντούν σε μεγάλες εκκρεμότητες που μας έρχονται από το χθες. Συμπερασματικά,
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, fairly guarded; reads as contempt; style: authoritative, monologue; average recording, quiet background; genuineness 2.6/6; vocal-burst blend 1.3/10; 11.2s, EL.
greece_olomeleia-20210713a_9327776_9339008 · in -24.2 dBFS · gain +4.2 dB · eurospeech-00790
(emotional numbness, pride, triumph · normally alert, slightly relaxed, some disfluency, monologue) αυτή τη στιγμή ξεδιπλώνεται στη χώρα ένα πρόγραμμα δημοσίων επενδύσεων το οποίο αγγίζει τα 13 δισεκατομμύρια ευρώ
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, normal breath; affect is neutral, slightly dominant, fairly guarded; reads as emotional numbness, pride, triumph; style: monologue, authoritative; average recording, quiet background; genuineness 1.8/6; vocal-burst blend 1.5/10; 10.5s, EL.
greece_olomeleia-20210713a_9339008_9349481 · in -23.8 dBFS · gain +3.8 dB · eurospeech-00790
Concentration ↓  /  Prideidentity +0.44 emotion 190 %   c-eurospeech-AB2 · #7

This chain comes from the two-sided rule: it only counts if both emotions move — Concentration down and Pride up — by at least 0.25 each.

The chain starts with Pride clearly present — 0.67, higher than 67 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.31.

At the same time Concentration goes the other way, from 0.96 (higher than 96 % of clips in this corpus) to 0.68 (higher than 68 % of clips in this corpus), a change of -0.28. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.23, then +0.02, then +0.05 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.47 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.49 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.47, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 69 s · da · eurospeech

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.468 before conversion and 0.906 after — it rose by 0.438. Neighbour-to-neighbour the worst pair went 0.463 → 0.900. (The earlier render, with segment 1 left raw, scores 0.796 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Pride moved +0.308 in the original and +0.586 after conversion — 190 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Concentration, -0.277 became -0.552.

Quality. Mean predicted overall quality across the segments went 3.31 → 3.54 (+0.23) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.468 → 0.906 +0.438identity cos neighbours 0.463 → 0.900d_b rescored +0.308 → +0.586d_a rescored -0.277 → -0.552d_a mined -0.276d_b mined 0.308min_cos_consec (site) 0.4856min_cos_anchor (site) 0.4702dataset eurospeechlang daspeaker denmark_20231M082_2024-04-total 68.1schain gain +2.3 dBseam step 0.5 dBcrossfades 150/150/150 ms
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: a middle-aged masculine voice · neutral-toned, average recording, quiet background, normally alert, light breath
(concentration, disappointment, disgust · measured, slightly relaxed, fairly steady, monologue) Men jeg synes alligevel, det spidser en lillebitte smule til her, for statsministeren har flere gange sagt, og det er også det, ordføreren refererer, at hvis man fra færøsk side ønsker en forandring, må der komme et ønske fra færøsk side om det. Det synes jeg jo i og for sig er en åben melding på det.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, disappointment, disgust; style: monologue, didactic; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 0.5/10; 18.5s, DA.
denmark_20231M082_2024-04-19_0900_11288896_11307424 · in -24.3 dBFS · gain +4.3 dB · eurospeech-00512
(concentration, pride · measured, slightly relaxed, moderately variable, monologue) Så træder vi vande og træder måske hinanden lidt over tæerne, når vi spørger, hvad ønsket så er med det her. Jeg har lyst til at sige til ordføreren, at jeg også synes, at det skal være legitimt, at jeg synes rigtig godt om rigsfællesskabet,
full caption & clip details
An elderly masculine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, pride; style: monologue, storytelling; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 3.7/10; 16.6s, DA.
denmark_20231M082_2024-04-19_0900_11307424_11324047 · in -24.0 dBFS · gain +4.0 dB · eurospeech-00512
(contempt, disgust, sourness · measured, slightly relaxed, fairly steady, didactic) for det gør jeg. Derfor er mit spørgsmål også til ordføreren: Hvad er det ordføreren tænker der skal komme, (low mumble) eller vil der komme noget fra færøsk side? Det er der også flere andre ordførere der har spurgt om her i salen. Hvad kan vi forvente?
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as contempt, disgust, sourness; style: didactic, monologue; average recording, quiet background; genuineness 2.0/6; vocal-burst blend 0.0/10; 14.6s, DA.
denmark_20231M082_2024-04-19_0900_11324047_11338671 · in -24.4 dBFS · gain +4.4 dB · eurospeech-00512
(pride, triumph, shame · normal-paced, neutral tension, moderately variable, casual) (ahem) (low mumble) Jeg tror først og fremmest, at vi kan forvente, at diskussionen fortsætter noget tid endnu. (ahem) Ellers noterede jeg mig, at (ahem) den færøske lagmand i går sagde, at partierne bliver nødt til at sætte sig ned ved et bord og finde ud af, hvad det er, man konkret vil,
full caption & clip details
An elderly somewhat feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, slightly dark, fairly smooth, thin; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as pride, triumph, shame; style: casual; average recording, quiet background; genuineness 5.1/6; vocal-burst blend 5.2/10; 19.0s, DA.
denmark_20231M082_2024-04-19_0900_11338671_11357632 · in -24.2 dBFS · gain +4.2 dB · eurospeech-00512
Thankfulness Gratitude ↓  /  Contemptidentity −0.02 emotion 187 %   c-eurospeech-AB2 · #8

This chain comes from the two-sided rule: it only counts if both emotions move — Thankfulness Gratitude down and Contempt up — by at least 0.25 each.

The chain starts with Contempt clearly present — 0.72, higher than 72 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.27.

At the same time Thankfulness Gratitude goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.66 (higher than 66 % of clips in this corpus), a change of -0.32. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are -0.11, then +0.19, then +0.08, then +0.11 — not a clean run: step 1 moves back the other way by 0.11 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.88 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.92 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.88. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 77 s · hr · eurospeech

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.888 before conversion and 0.867 after — it fell by 0.021. Neighbour-to-neighbour the worst pair went 0.911 → 0.848. (The earlier render, with segment 1 left raw, scores 0.699 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contempt moved +0.269 in the original and +0.502 after conversion — 187 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Thankfulness Gratitude, -0.318 became -0.155.

Quality. Mean predicted overall quality across the segments went 3.18 → 3.46 (+0.29) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.888 → 0.867 -0.021identity cos neighbours 0.911 → 0.848d_b rescored +0.269 → +0.502d_a rescored -0.318 → -0.155d_a mined -0.318d_b mined 0.271min_cos_consec (site) 0.9196min_cos_anchor (site) 0.8819dataset eurospeechlang hrspeaker croatia_20210624091825-219total 75.8schain gain +3.2 dBseam step 0.8 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: an adult feminine voice · neutral-bright, thin, average recording, quiet background, neutral tension, moderately variable, some disfluency, light breath
(thankfulness gratitude, shame, pride · normal-paced, normally alert, somewhat unclear, casual) I tada smo rekli da nam je (low mumble) manje važno, dakle tko je koga predložio i da ćemo se zaista voditi objektivnim (low mumble) kriterijima tko je od kandidata i kandidatkinja zaista najstručniji, tko ima
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, thin; somewhat unclear, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as thankfulness gratitude, shame, pride; style: casual, dramatic; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 7.9/10; 14.7s, HR.
croatia_20210624091825-21928_20363728_20378384 · in -14.1 dBFS · gain -5.9 dB · eurospeech-01532
(pride, shame, triumph · normal-paced, normally alert, average clarity, monologue) (ahem) iza sebe najbolje iskustvo i tko u svom programu je pokazao širinu i shvaćanje ključnih problema pravosuđa. Ja mislim da svatko od vas pojedinačno bez obzira na to kako ćete glasati,
full caption & clip details
An elderly feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as pride, shame, triumph; style: monologue, dramatic; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 5.8/10; 16.0s, HR.
croatia_20210624091825-21928_20378384_20394400 · in -15.6 dBFS · gain -4.4 dB · eurospeech-01532
(pride, thankfulness gratitude · normal-paced, normally alert, average clarity, monologue) ako je pristupio čitanju programa na taj način da je vidio je program prof. Đurđević kao i njezino odgovaranje na naša pitanja na saslušanju
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, thin; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as pride, thankfulness gratitude; style: monologue, dramatic; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 6.2/10; 10.7s, HR.
croatia_20210624091825-21928_20394400_20405088 · in -15.2 dBFS · gain -4.8 dB · eurospeech-01532
(triumph, bitterness, interest · normal-paced, normally alert, somewhat unclear, monologue) (ahem) pokazuje jednostavno veliku razliku u odnosu na sve druge kandidate i po stručnosti i po kvaliteti, ali napominjem i po integritetu koji je pokazala. Na žalost (ahem) ono što je danas zapravo ispalo u ovoj raspravi je da nismo
full caption & clip details
An elderly feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, slightly rough, thin; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as triumph, bitterness, interest; style: monologue, casual; average recording, quiet background; genuineness 4.5/6; vocal-burst blend 6.3/10; 19.5s, HR.
croatia_20210624091825-21928_20405088_20424591 · in -15.6 dBFS · gain -4.4 dB · eurospeech-01532
(contempt, shame, concentration · brisk, energised, very clear, dramatic) kao saborski zastupnici dorasli toj ideji da možemo kao političko tijelo, kao netko tko je politički reprezent građana u Hrvatskoj birati po nepolitičkoj liniji, odnosno
full caption & clip details
An adult feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; very clear, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as contempt, shame, concentration; style: dramatic, cartoonish; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 5.9/10; 15.7s, HR.
croatia_20210624091825-21928_20424591_20440320 · in -15.1 dBFS · gain -5.0 dB · eurospeech-01532
Interest ↓  /  Shameidentity +0.03 emotion 148 %   c-eurospeech-AB2 · #9

This chain comes from the two-sided rule: it only counts if both emotions move — Interest down and Shame up — by at least 0.25 each.

The chain starts with Shame clearly present — 0.62, higher than 62 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.35.

At the same time Interest goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.67 (higher than 67 % of clips in this corpus), a change of -0.30. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.20, then -0.07, then +0.22 — not a clean run: step 2 moves back the other way by 0.07 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.88 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.89 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.88. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 53 s · en · eurospeech

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.875 before conversion and 0.906 after — it rose by 0.030. Neighbour-to-neighbour the worst pair went 0.881 → 0.914. (The earlier render, with segment 1 left raw, scores 0.810 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Shame moved +0.349 in the original and +0.516 after conversion — 148 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Interest, -0.303 became -0.274.

Quality. Mean predicted overall quality across the segments went 3.02 → 3.29 (+0.26) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.875 → 0.906 +0.030identity cos neighbours 0.881 → 0.914d_b rescored +0.349 → +0.516d_a rescored -0.303 → -0.274d_a mined -0.302d_b mined 0.350min_cos_consec (site) 0.8947min_cos_anchor (site) 0.8777dataset eurospeechlang enspeaker uk_uk_32_06112018total 51.9schain gain +2.8 dBseam step 0.3 dBcrossfades 150/150/150 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: an adult masculine voice · neutral-toned, neutral-bright, normally alert, fairly steady, light breath
(interest, fear, triumph · measured, neutral tension, some disfluency, authoritative) nobody had heard of post-traumatic stress disorder, but today the issue is not just what we can do for our veterans returning from the frontline, but how we can prioritise the mental health of everyone. One hundred years ago,
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as interest, fear, triumph; style: authoritative, conversational; good recording, no background noise; genuineness 1.9/6; vocal-burst blend 3.1/10; 13.7s, EN.
uk_uk_32_06112018_19669568_19683232 · in -23.0 dBFS · gain +3.0 dB · eurospeech-01002
(sadness, distress, helplessness · measured, neutral tension, some disfluency, monologue) people from all over the world fought and died to protect our country, and today we need to remember the debt we owe people who were not born here but who helped to make our country what it is today. One hundred years ago, the first world war changed
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as sadness, distress, helplessness; style: monologue, formal; average recording, quiet background; genuineness 1.5/6; vocal-burst blend 1.2/10; 15.3s, EN.
uk_uk_32_06112018_19683232_19698544 · in -22.1 dBFS · gain +2.0 dB · eurospeech-01002
(concentration · measured, slightly relaxed, little disfluency, formal) the role of the state; Government took action on food, rents and wages. That links to one of the central arguments in our public life today: what Governments should and should not do in the 21st century.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, little disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as concentration; style: formal, monologue; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 0.8/10; 12.9s, EN.
uk_uk_32_06112018_19698544_19711424 · in -23.5 dBFS · gain +3.5 dB · eurospeech-01002
(shame, thankfulness gratitude, affection · normal-paced, neutral tension, some disfluency, conversational) (low mumble) that in due course we will remember not only those who fell in the service of our country in the first world war, but those who have fallen more recently.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, normal-paced, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is mildly positive, slightly dominant, neutral openness; reads as shame, thankfulness gratitude, affection; style: conversational, casual; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 1.3/10; 10.7s, EN.
uk_uk_32_06112018_19730464_19741152 · in -22.8 dBFS · gain +2.8 dB · eurospeech-01002
Concentration ↓  /  Prideidentity +0.61 emotion 157 %   c-eurospeech-AB2 · #10

This chain comes from the two-sided rule: it only counts if both emotions move — Concentration down and Pride up — by at least 0.25 each.

The chain starts with Pride clearly present — 0.67, higher than 67 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.27.

At the same time Concentration goes the other way, from 0.94 (higher than 94 % of clips in this corpus) to 0.68 (higher than 68 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.06, then +0.21 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.11 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.08 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.11, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 48 s · lv · eurospeech

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.106 before conversion and 0.715 after — it rose by 0.609. Neighbour-to-neighbour the worst pair went 0.075 → 0.721. (The earlier render, with segment 1 left raw, scores 0.711 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Pride moved +0.267 in the original and +0.419 after conversion — 157 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Concentration, -0.258 became -0.469.

Quality. Mean predicted overall quality across the segments went 2.70 → 3.32 (+0.62) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.106 → 0.715 +0.609identity cos neighbours 0.075 → 0.721d_b rescored +0.267 → +0.419d_a rescored -0.258 → -0.469d_a mined -0.259d_b mined 0.266min_cos_consec (site) 0.0751min_cos_anchor (site) 0.1099dataset eurospeechlang lvspeaker latvia_20130214103407total 47.3schain gain +2.6 dBseam step 0.1 dBcrossfades 150/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: a child feminine voice · slightly cool, average recording, quiet background
(concentration · brisk, energised, neutral tension, dramatic) bet administratīvi teritoriālās reformas rezultātā tas tika palielināts līdz 13 deputātiem. Šobrīd šis priekšlikums paredz atgriezties pie (ahem) sākotnējā deputātu skaita - tātad pie 9 deputātiem. (ahem)
full caption & clip details
A child feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, slightly guarded; reads as concentration; style: dramatic, cartoonish; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 4.7/10; 15.0s, LV.
latvia_20130214103407_905040_920040 · in -22.2 dBFS · gain +2.2 dB · eurospeech-01993
(contempt, disgust, disappointment · fast, energised, neutral tension, cartoonish) Komisijā diskusijās piedalījās arī Latvijas Pašvaldību savienība, kas pauda viedokli, ka lielākā daļa no mazajām pašvaldībām, kurās iedzīvotāju skaits ir mazāks par 5 tūkstošiem, ir piekritušas šādām izmaiņām, jo uzskata, ka šādas izmaiņas padarīs pašvaldības darbu efektīvāku.
full caption & clip details
A child feminine voice; delivery is energised, fast, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, thin; clear, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as contempt, disgust, disappointment; style: cartoonish, dramatic; average recording, quiet background; genuineness 2.5/6; vocal-burst blend 5.3/10; 16.2s, LV.
latvia_20130214103407_920040_936240 · in -22.4 dBFS · gain +2.4 dB · eurospeech-01993
(pride, thankfulness gratitude, shame · measured, subdued, slightly relaxed, monologue) Ļoti cienījamā Saeimas priekšsēdētāja! Dāmas un kungi! Protams, par šo jautājumu, kas saistīts ar deputātu skaita samazināšanu pašvaldībās, mēs esam diskutējuši jau ļoti ilgi
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is slightly cool, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, fairly guarded; reads as pride, thankfulness gratitude, shame; style: monologue, whispered; average recording, quiet background; genuineness 2.3/6; vocal-burst blend 0.7/10; 16.4s, LV.
latvia_20130214103407_952567_969008 · in -25.6 dBFS · gain +5.5 dB · eurospeech-01993
Thankfulness Gratitude ↓  /  Interestidentity +0.02 emotion 89 %   c-eurospeech-AB2 · #11

This chain comes from the two-sided rule: it only counts if both emotions move — Thankfulness Gratitude down and Interest up — by at least 0.25 each.

The chain starts with Interest clearly present — 0.73, higher than 73 % of clips in this corpus — and ends with it at the very top of the corpus at 0.99, higher than 99 % of clips in this corpus. That is a total rise of 0.26.

At the same time Thankfulness Gratitude goes the other way, from 0.95 (higher than 95 % of clips in this corpus) to 0.55 (higher than 55 % of clips in this corpus), a change of -0.41. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.02, then +0.20, then +0.04 — a plateau around step 1, where it barely moves.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.88 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.88 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.88. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 60 s · en · eurospeech

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.863 before conversion and 0.885 after — it rose by 0.022. Neighbour-to-neighbour the worst pair went 0.831 → 0.888. (The earlier render, with segment 1 left raw, scores 0.679 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Interest moved +0.255 in the original and +0.227 after conversion — 89 % of the delta retained, which is most of it. On the other named axis, Thankfulness Gratitude, -0.406 became -0.365.

Quality. Mean predicted overall quality across the segments went 3.03 → 3.23 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.863 → 0.885 +0.022identity cos neighbours 0.831 → 0.888d_b rescored +0.255 → +0.227d_a rescored -0.406 → -0.365d_a mined -0.406d_b mined 0.256min_cos_consec (site) 0.8784min_cos_anchor (site) 0.8763dataset eurospeechlang enspeaker uk_uk_4_30102024total 59.3schain gain +4.2 dBseam step 0.7 dBcrossfades 150/150/150 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: an elderly masculine voice · neutral-toned, neutral-bright, full, measured, normally alert, wide pitch range
(thankfulness gratitude, relief, contentment · neutral tension, moderately variable, some disfluency, conversational) I am glad that the review will look again at getting rid of the cliff edge for carer’s allowance (ahem) and the earnings limit, I hope she and her colleagues will consider a broader review to give family carers the support they deserve.
full caption & clip details
An elderly masculine voice; delivery is normally alert, measured, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, full; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as thankfulness gratitude, relief, contentment; style: conversational, storytelling; good recording, quiet background; genuineness 2.7/6; vocal-burst blend 3.8/10; 16.8s, EN.
uk_uk_4_30102024_11137216_11153968 · in -25.3 dBFS · gain +5.3 dB · eurospeech-01049
(disappointment, bitterness, anger · slightly relaxed, fairly steady, almost no disfluency, newsreading) The Conservative cost of living crisis has hit family carers particularly hard, but they are not alone. Practically no one has been left untouched by rising bills, higher mortgage payments and soaring food prices.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; clear, almost no disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as disappointment, bitterness, anger; style: newsreading, storytelling; average recording, no background noise; genuineness 0.3/6; vocal-burst blend 0.3/10; 18.6s, EN.
uk_uk_4_30102024_11153968_11172528 · in -24.8 dBFS · gain +4.8 dB · eurospeech-01049
(interest, concentration, contemplation · neutral tension, moderately variable, some disfluency, dramatic) Our small businesses and the self-employed have experienced a crisis of their own, having had to deal with the pandemic and now spiking energy prices and other input costs.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, full; average clarity, some disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, slightly guarded; reads as interest, concentration, contemplation; style: dramatic, storytelling; average recording, quiet background; genuineness 1.7/6; vocal-burst blend 0.9/10; 11.4s, EN.
uk_uk_4_30102024_11172528_11183936 · in -25.9 dBFS · gain +5.9 dB · eurospeech-01049
(interest, anger, bitterness · neutral tension, moderately variable, some disfluency, dramatic) The Conservatives only added to that pain by hitting struggling families with stealth tax rises, by betraying pensioners when they broke the triple lock, and by raising national insurance for our small businesses. I welcome
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, full; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, slightly guarded; reads as interest, anger, bitterness; style: dramatic, storytelling; good recording, quiet background; genuineness 2.6/6; vocal-burst blend 2.6/10; 13.2s, EN.
uk_uk_4_30102024_11183936_11197088 · in -25.4 dBFS · gain +5.4 dB · eurospeech-01049
Concentration ↓  /  Reliefidentity +0.63 emotion 95 %   c-eurospeech-AB2 · #12

This chain comes from the two-sided rule: it only counts if both emotions move — Concentration down and Relief up — by at least 0.25 each.

The chain starts with Relief around average — 0.53, higher than 53 % of clips in this corpus — and ends with it at the very top of the corpus at 0.98, higher than 98 % of clips in this corpus. That is a total rise of 0.45.

At the same time Concentration goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.52 (higher than 52 % of clips in this corpus), a change of -0.45. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.23, then +0.06, then +0.10, then +0.06 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.12 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.11 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.12, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 74 s · fr · eurospeech

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.104 before conversion and 0.736 after — it rose by 0.632. Neighbour-to-neighbour the worst pair went 0.110 → 0.746. (The earlier render, with segment 1 left raw, scores 0.607 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Relief moved +0.449 in the original and +0.427 after conversion — 95 % of the delta retained, which is essentially all of it. On the other named axis, Concentration, -0.447 became -0.442.

Quality. Mean predicted overall quality across the segments went 3.05 → 3.29 (+0.24) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.104 → 0.736 +0.632identity cos neighbours 0.110 → 0.746d_b rescored +0.449 → +0.427d_a rescored -0.447 → -0.442d_a mined -0.447d_b mined 0.449min_cos_consec (site) 0.1140min_cos_anchor (site) 0.1159dataset eurospeechlang frspeaker france_france_senate_33020total 73.0schain gain +1.0 dBseam step 0.9 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: an adult masculine voice · average recording, quiet background
(concentration, pride, disappointment · brisk, energised, neutral tension, dramatic) A dit qu'il regrettait la suppression de l'index, a dit qu'il considérait que l'emploi des seniors ne progressera qu'avec le dialogue social et qu'Supprimons l'index. L'Assemblée nationale, avait pris le chemin inverse et Lise ça notamment parce qu'on renvoie la définition des critères et la possibilité de leur adaptation
full caption & clip details
An adult masculine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as concentration, pride, disappointment; style: dramatic, ranting; average recording, quiet background; genuineness 2.6/6; vocal-burst blend 4.3/10; 14.1s, FR.
france_france_senate_3302062_6404941d5ef71_12490736_12504848 · in -28.1 dBFS · gain +8.1 dB · eurospeech-01174
(concentration, sourness, bitterness · fast, energised, neutral tension, dramatic) au dialogue social interprofessionnel pour le décret. Petite question des critères et au dialogue social de branche pour l'adaptation. C'est pour ça que nous avons un index qui renvoie à à ce rôle des partenaires sociaux pour définir les bons indicateurs et la bonne mise en oeuvre. Donc je m'arrête là mais évidemment pour ces 2 raisons complémentaires en plus de ce que j'ai dit, l'avis est défavorable.
full caption & clip details
A young adult masculine voice; delivery is energised, fast, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as concentration, sourness, bitterness; style: dramatic, ranting; average recording, quiet background; genuineness 2.3/6; vocal-burst blend 6.8/10; 16.4s, FR.
france_france_senate_3302062_6404941d5ef71_12504848_12521264 · in -28.4 dBFS · gain +8.4 dB · eurospeech-01174
(sourness, disgust, thankfulness gratitude · slow, very low-energy, relaxed, monologue) Bien donc nous avons un double avis défavorable de la Commission et (low mumble) du ministre-sur ces amendements de suppression avant de les soumettre aux voix j'ai des demandes d'explications de vote. Madame Carlotti.
full caption & clip details
An elderly masculine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is neutral-toned, slightly dark, rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as sourness, disgust, thankfulness gratitude; style: monologue, whispered; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 0.8/10; 17.9s, FR.
france_france_senate_3302062_6404941d5ef71_12521264_12539159 · in -28.9 dBFS · gain +8.9 dB · eurospeech-01174
(affection, thankfulness gratitude, relief · brisk, energised, neutral tension, cartoonish) Monsieur le Président. Moi je souhaite soutenir les amendements de suppression de cet article de car nous ne devrions même pas en discuter, ça a été très bien dit par ma collègue tout à l'heure. Hé oui d'après la vie.
full caption & clip details
An elderly feminine voice; delivery is energised, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; average clarity, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, vulnerable; reads as affection, thankfulness gratitude, relief; style: cartoonish, dramatic; average recording, quiet background; genuineness 3.0/6; vocal-burst blend 3.9/10; 15.0s, FR.
france_france_senate_3302062_6404941d5ef71_12539159_12554176 · in -25.9 dBFS · gain +5.8 dB · eurospeech-01174
(relief, disappointment, emotional numbness · fast, energised, neutral tension, ranting) Du Conseil d'État, il s'agirait d'un cavalier social qui pourrait même être inconstitutionnelle. Alors tout à l'heure, la droite n'a pas voulu que cet article soit renvoyé en commission mais c'est vraiment fort dommage
full caption & clip details
A young adult feminine voice; delivery is energised, fast, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, thin; clear, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, guarded; reads as relief, disappointment, emotional numbness; style: ranting, dramatic; average recording, quiet background; genuineness 2.4/6; vocal-burst blend 4.6/10; 10.4s, FR.
france_france_senate_3302062_6404941d5ef71_12554176_12564560 · in -26.6 dBFS · gain +6.5 dB · eurospeech-01174
Relief ↓  /  Emotional Numbnessidentity +0.39 emotion 22 %   c-eurospeech-AB2 · #13

This chain comes from the two-sided rule: it only counts if both emotions move — Relief down and Emotional Numbness up — by at least 0.25 each.

The chain starts with Emotional Numbness clearly present — 0.65, higher than 65 % of clips in this corpus — and ends with it at the very top of the corpus at 0.91, higher than 91 % of clips in this corpus. That is a total rise of 0.27.

At the same time Relief goes the other way, from 0.90 (higher than 90 % of clips in this corpus) to 0.50 (right about the corpus median), a change of -0.40. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are -0.07, then +0.20, then +0.14 — not a clean run: step 1 moves back the other way by 0.07 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.30 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.31 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.30, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 52 s · fr · eurospeech

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.324 before conversion and 0.719 after — it rose by 0.395. Neighbour-to-neighbour the worst pair went 0.300 → 0.686. (The earlier render, with segment 1 left raw, scores 0.541 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Emotional Numbness moved +0.266 in the original and +0.060 after conversion — 22 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Relief, -0.401 became -0.366.

Quality. Mean predicted overall quality across the segments went 3.13 → 3.30 (+0.18) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.324 → 0.719 +0.395identity cos neighbours 0.300 → 0.686d_b rescored +0.266 → +0.060d_a rescored -0.401 → -0.366d_a mined -0.401d_b mined 0.265min_cos_consec (site) 0.3120min_cos_anchor (site) 0.2953dataset eurospeechlang frspeaker france_france_senate_29830total 50.7schain gain -0.1 dBseam step 0.8 dBcrossfades 150/150/150 ms
Script — 4 chunks, 0 with a non-speech sound
Unchanged across all 4 clips: a middle-aged masculine voice · neutral-toned, slightly rough, balanced body, quiet background, measured, slightly relaxed
(relief · normally alert, steady, some disfluency, monologue) Et comme l'inflation pousse les taux d'intérêt à la hausse, nous allons redécouvrir que s'endetter a un coût, que nous ne pouvons plus nous permettre sans renoncer à des priorités essentielles.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, slightly guarded; reads as relief; style: monologue; average recording, quiet background; genuineness 1.5/6; vocal-burst blend 1.3/10; 10.5s, FR.
france_france_senate_2983040_62ebd257ce18c_3531184_3541712 · in -28.1 dBFS · gain +8.1 dB · eurospeech-01153
(sadness, anger, triumph · subdued, steady, frequent disfluency, monologue) Le « quoi qu'il en coûte » nous a permis de tenir bon et de surmonter la crise sanitaire, économique et sociale. Il était nécessaire. Mais il faut désormais remettre de l'ordre dans nos comptes, pour nous préparer aux défis qui s'annoncent. La parole
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as sadness, anger, triumph; style: monologue, whispered; average recording, quiet background; genuineness 1.4/6; vocal-burst blend 0.7/10; 16.8s, FR.
france_france_senate_2983040_62ebd257ce18c_3541712_3558560 · in -28.1 dBFS · gain +8.1 dB · eurospeech-01153
(normally alert, fairly steady, some disfluency, monologue) la CMP qui s'est réunie hier soir est parvenue à un très bon accord, sur un texte particulièrement important, qui vise la protection du pouvoir d'achat de nos compatriotes.
full caption & clip details
A child masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is positive, neutral stance, slightly guarded; no dominant emotion; style: monologue, narration; good recording, quiet background; genuineness 0.7/6; vocal-burst blend 0.0/10; 10.5s, FR.
france_france_senate_2983040_62ebd257ce18c_3585488_3595969 · in -30.2 dBFS · gain +10.2 dB · eurospeech-01153
(emotional numbness, concentration · normally alert, fairly steady, some disfluency, monologue) Les débats ont été longs et intenses et ils ont permis, me semble-t-il, à tous les groupes de défendre leurs propositions pour soutenir les Français marqués par la hausse de l'inflation. Pour ce qui concerne le groupe Les Républicains, notre ligne a été très claire
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, concentration; style: monologue, didactic; average recording, quiet background; genuineness 1.8/6; vocal-burst blend 0.0/10; 13.4s, FR.
france_france_senate_2983040_62ebd257ce18c_3595969_3609376 · in -29.1 dBFS · gain +9.1 dB · eurospeech-01153
Triumph ↓  /  Contemptidentity +0.64 emotion 24 %   c-eurospeech-AB2 · #14

This chain comes from the two-sided rule: it only counts if both emotions move — Triumph down and Contempt up — by at least 0.25 each.

The chain starts with Contempt clearly present — 0.70, higher than 70 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.27.

At the same time Triumph goes the other way, from 0.93 (higher than 93 % of clips in this corpus) to 0.56 (higher than 56 % of clips in this corpus), a change of -0.36. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.25, then -0.05, then +0.07 — not a clean run: step 2 moves back the other way by 0.05 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.10 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.15 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.10, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 66 s · lv · eurospeech

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.105 before conversion and 0.743 after — it rose by 0.638. Neighbour-to-neighbour the worst pair went 0.151 → 0.742. (The earlier render, with segment 1 left raw, scores 0.337 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Contempt moved +0.274 in the original and +0.066 after conversion — 24 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Triumph, -0.364 became -0.524.

Quality. Mean predicted overall quality across the segments went 2.87 → 3.28 (+0.41) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.105 → 0.743 +0.638identity cos neighbours 0.151 → 0.742d_b rescored +0.274 → +0.066d_a rescored -0.364 → -0.524d_a mined -0.364d_b mined 0.273min_cos_consec (site) 0.1487min_cos_anchor (site) 0.0955dataset eurospeechlang lvspeaker latvia_20210804172801total 64.9schain gain +4.4 dBseam step 3.5 dBcrossfades 150/150/150 ms
Script — 4 chunks, 1 with a non-speech sound
Unchanged across all 4 clips: a young adult masculine voice · balanced body, average recording, quiet background, measured, slightly relaxed
(triumph, pride · normally alert, fairly steady, frequent disfluency, monologue) Rezultātā komisijā šis likumprojekts tika atbalstīts pirmajā lasījumā. Komisija aicina to atbalstīt arī pirmajā lasījumā parlamentā. Sēdes vadītāja. (ahem) Paldies par ziņojumu. Sākam debates. Vārds deputātam Aldim Gobzemam. Lūdzu!
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, dark, fairly smooth, balanced body; slurred, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as triumph, pride; style: monologue; average recording, quiet background; genuineness 3.9/6; vocal-burst blend 1.6/10; 15.9s, LV.
latvia_20210804172801_1322608_1338464 · in -31.8 dBFS · gain +11.8 dB · eurospeech-02044
(anger, shame, emotional numbness · subdued, fairly steady, frequent disfluency, monologue) Latvijas sabiedrība! Es paturpināšu par to visu. Tātad obligāta vakcinācija. Šie nelieši neviens ne personīgi, ne kā citādi neatbildēs
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is slightly cool, slightly dark, slightly rough, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, fairly guarded; reads as anger, shame, emotional numbness; style: monologue, authoritative; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 0.5/10; 17.5s, LV.
latvia_20210804172801_1338464_1356000 · in -29.6 dBFS · gain +9.6 dB · eurospeech-02044
(anger, shame, impatience and irritability · subdued, steady, frequent disfluency, didactic) par to, ja kaut kas notiks manai mātei vai maniem bērniem, vai man, vai jums, dārgie skatītāji, dārgie Latvijas Republikas pilsoņi. Viņi balso.
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, steady; timbre is neutral-toned, slightly dark, slightly rough, balanced body; slurred, frequent disfluency, fairly narrow pitch, normal breath; affect is neutral, neutral stance, fairly guarded; reads as anger, shame, impatience and irritability; style: didactic, monologue; average recording, quiet background; genuineness 3.3/6; vocal-burst blend 0.6/10; 16.2s, LV.
latvia_20210804172801_1356000_1372240 · in -29.4 dBFS · gain +9.4 dB · eurospeech-02044
(contempt, anger, disappointment · subdued, fairly steady, some disfluency, authoritative) Viņi balso, neuzņemoties nekādu atbildību, viņi balso par to, ka jūs ir obligāti jāvakcinē un jāievieš kodi kā pirms kara Vācijā 1933. gadā,
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, slightly relaxed, fairly steady; timbre is slightly cool, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, fairly narrow pitch, normal breath; affect is neutral, slightly dominant, fairly guarded; reads as contempt, anger, disappointment; style: authoritative, monologue; average recording, quiet background; genuineness 2.4/6; vocal-burst blend 1.6/10; 15.9s, LV.
latvia_20210804172801_1372240_1388112 · in -26.0 dBFS · gain +6.0 dB · eurospeech-02044
Contempt ↓  /  Thankfulness Gratitudeidentity +0.01 emotion 49 %   c-eurospeech-AB2 · #15

This chain comes from the two-sided rule: it only counts if both emotions move — Contempt down and Thankfulness Gratitude up — by at least 0.25 each.

The chain starts with Thankfulness Gratitude around average — 0.55, higher than 55 % of clips in this corpus — and ends with it strongly present at 0.89, higher than 89 % of clips in this corpus. That is a total rise of 0.34.

At the same time Contempt goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.61 (higher than 61 % of clips in this corpus), a change of -0.36. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.25, then +0.10 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.87 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.92 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.87. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 37 s · lt · eurospeech

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.852 before conversion and 0.859 after — it rose by 0.007. Neighbour-to-neighbour the worst pair went 0.908 → 0.829. (The earlier render, with segment 1 left raw, scores 0.620 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Thankfulness Gratitude moved +0.344 in the original and +0.168 after conversion — 49 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Contempt, -0.357 became -0.409.

Quality. Mean predicted overall quality across the segments went 3.04 → 3.45 (+0.40) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.852 → 0.859 +0.007identity cos neighbours 0.908 → 0.829d_b rescored +0.344 → +0.168d_a rescored -0.357 → -0.409d_a mined -0.358d_b mined 0.344min_cos_consec (site) 0.9174min_cos_anchor (site) 0.8722dataset eurospeechlang ltspeaker lithuania_lithuania_2_2404total 36.3schain gain +1.1 dBseam step 0.1 dBcrossfades 150/150 ms
Script — 3 chunks, 0 with a non-speech sound
Unchanged across all 3 clips: a middle-aged masculine voice · balanced body, quiet background, normally alert, fairly steady, some disfluency
(contempt, triumph · measured, neutral tension, average clarity, cartoonish) Aš tiesiog pasidomėjau, kaip yra kituose nacionaliniuose parkuose, draustiniuose, rezervatuose ir t.t. Jūs įsivaizduojate, aš gavau atsakymą, kad faktiškai visose
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, neutral tension, fairly steady; timbre is slightly cool, slightly dark, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, fairly guarded; reads as contempt, triumph; style: cartoonish, authoritative; below-average recording, quiet background; genuineness 3.1/6; vocal-burst blend 1.5/10; 13.4s, LT.
lithuania_lithuania_2_24042001_5654528_5667920 · in -11.0 dBFS · gain -9.0 dB · eurospeech-01931
(triumph, relief · brisk, slightly relaxed, somewhat unclear, authoritative) Problema yra ta, kad tie, kurie sudarinėjo tų saugomų teritorijų ribas, tikrai nepažiūrėjo į tai, kas vyksta. Imkime, pavyzdžiui, Čepkelių raistą, mums visiems jis žinomas, bet iš esmės tai yra kaimas.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as triumph, relief; style: authoritative, monologue; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 4.5/10; 13.0s, LT.
lithuania_lithuania_2_24042001_5693712_5706752 · in -10.5 dBFS · gain -9.5 dB · eurospeech-01931
(brisk, slightly relaxed, average clarity, didactic) Pagal įstatymą kaimo gyventojai yra nelegalai, juos reikia iškelti, padaryti tikrą rezervatą ir susodinti vaškines figūras. Taip turi būti.
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; no dominant emotion; style: didactic, authoritative; average recording, quiet background; genuineness 3.1/6; vocal-burst blend 1.4/10; 10.3s, LT.
lithuania_lithuania_2_24042001_5706752_5717017 · in -10.6 dBFS · gain -9.4 dB · eurospeech-01931
Shame ↓  /  Concentrationidentity −0.02 emotion 136 %   c-eurospeech-AB2 · #16

This chain comes from the two-sided rule: it only counts if both emotions move — Shame down and Concentration up — by at least 0.25 each.

The chain starts with Concentration around average — 0.55, higher than 55 % of clips in this corpus — and ends with it at the very top of the corpus at 0.93, higher than 93 % of clips in this corpus. That is a total rise of 0.37.

At the same time Shame goes the other way, from 0.97 (higher than 97 % of clips in this corpus) to 0.60 (higher than 60 % of clips in this corpus), a change of -0.37. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.12, then -0.16, then +0.24, then +0.18 — not a clean run: step 2 moves back the other way by 0.16 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.90 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.90 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.90), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 72 s · hr · eurospeech

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.891 before conversion and 0.874 after — it fell by 0.017. Neighbour-to-neighbour the worst pair went 0.891 → 0.833. (The earlier render, with segment 1 left raw, scores 0.649 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.371 in the original and +0.505 after conversion — 136 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Shame, -0.367 became -0.036.

Quality. Mean predicted overall quality across the segments went 3.00 → 3.27 (+0.27) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.891 → 0.874 -0.017identity cos neighbours 0.891 → 0.833d_b rescored +0.371 → +0.505d_a rescored -0.367 → -0.036d_a mined -0.367d_b mined 0.373min_cos_consec (site) 0.9007min_cos_anchor (site) 0.9007dataset eurospeechlang hrspeaker croatia_20081209173210-158total 70.3schain gain +0.3 dBseam step 1.2 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 0 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · neutral-toned, neutral-bright, balanced body, average recording, quiet background, normally alert, slightly relaxed, fairly steady
(shame, pride · brisk, almost no disfluency, moderate pitch range, authoritative) Već je iduće godine održan i prvi službeni turnir. O izgledu igrališta ne zna se mnogo ali postoji podatak da je početkom 30-tih godina duljina igrališta iznosila 4 tisuće 450 jardi.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, almost no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as shame, pride; style: authoritative, monologue; average recording, quiet background; genuineness 0.5/6; vocal-burst blend 2.7/10; 13.0s, HR.
croatia_20081209173210-15861_8976784_8989776 · in -27.7 dBFS · gain +7.7 dB · eurospeech-01267
(thankfulness gratitude, hope enthusiasm optimism, triumph · brisk, some disfluency, wide pitch range, authoritative) Po nekim izvorima prvi pokušaj osnivanja golf kluba u Zagrebu bio je 1925. godine ali postoje dokumenti prema kojima je Gradsko poglavarstvo tada odbilo dati pristanak na osnivanje kluba.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as thankfulness gratitude, hope enthusiasm optimism, triumph; style: authoritative, monologue; average recording, quiet background; genuineness 1.8/6; vocal-burst blend 3.7/10; 13.5s, HR.
croatia_20081209173210-15861_8989776_9003248 · in -26.1 dBFS · gain +6.1 dB · eurospeech-01267
(pride, relief, thankfulness gratitude · brisk, almost no disfluency, wide pitch range, authoritative) Ponovno osnivanje kluba uspješno je provedeno 1929.odine a utemeljitelji Golf kluba Zagreba bile su najviđenije osobe ondašnjeg Zagreba na čelu sa grofom Miroslavom Kulmerom.
full caption & clip details
A young adult masculine voice; delivery is normally alert, brisk, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, almost no disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, fairly guarded; reads as pride, relief, thankfulness gratitude; style: authoritative, monologue; average recording, quiet background; genuineness 0.9/6; vocal-burst blend 2.9/10; 12.4s, HR.
croatia_20081209173210-15861_9003248_9015696 · in -26.2 dBFS · gain +6.2 dB · eurospeech-01267
(pride, triumph, relief · normal-paced, some disfluency, moderate pitch range, authoritative) Osnivač kluba bio je i hotel Esplanada u čijim se prostorima nalazilo i sjedište kluba. 1931. godine klub ima 80-tak članova, a već 1932. klub broji oko 102 člana. Toliki broj članova za ondašnji Zagreb je bio doista respektabilan.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as pride, triumph, relief; style: authoritative, monologue; average recording, quiet background; genuineness 1.3/6; vocal-burst blend 3.7/10; 19.0s, HR.
croatia_20081209173210-15861_9015696_9034720 · in -27.4 dBFS · gain +7.4 dB · eurospeech-01267
(concentration · normal-paced, some disfluency, moderate pitch range, authoritative) Do početka 1931. godine sagrađeno je igralište od 9 polja. Igralište se svečano, je svečano otvoreno 12. lipnja 1931. godine uz nazočnost mnogih velikodostojnika.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as concentration; style: authoritative, monologue; average recording, quiet background; genuineness 1.1/6; vocal-burst blend 2.0/10; 13.2s, HR.
croatia_20081209173210-15861_9034720_9047920 · in -27.6 dBFS · gain +7.6 dB · eurospeech-01267
Thankfulness Gratitude ↓  /  Concentrationidentity +0.04 emotion 96 %   c-eurospeech-AB2 · #17

This chain comes from the two-sided rule: it only counts if both emotions move — Thankfulness Gratitude down and Concentration up — by at least 0.25 each.

The chain starts with Concentration clearly present — 0.70, higher than 70 % of clips in this corpus — and ends with it at the very top of the corpus at 1.00, virtually no clip in this corpus scores higher. That is a total rise of 0.29.

At the same time Thankfulness Gratitude goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.71 (higher than 71 % of clips in this corpus), a change of -0.27. Both halves had to happen for this chain to qualify.

It takes 4 clips to get there. Clip to clip the moves are +0.19, then +0.06, then +0.04 — most of the change happening immediately, then levelling off.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.90 against the first clip, where 1.00 would mean an identical voice. That is a strong match — almost certainly one person throughout. Neighbouring clips score at worst 0.94 against each other.

Voice consistency: these clips are separate recordings joined together. The measured match is tight (0.90), so any shift should be subtle — but you may still notice the voice change slightly from segment to segment. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

4 clips · 62 s · en · eurospeech

What was done to this chain. All 4 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.859 before conversion and 0.899 after — it rose by 0.039. Neighbour-to-neighbour the worst pair went 0.933 → 0.941. (The earlier render, with segment 1 left raw, scores 0.868 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.291 in the original and +0.281 after conversion — 96 % of the delta retained, which is essentially all of it. On the other named axis, Thankfulness Gratitude, -0.266 became -0.349.

Quality. Mean predicted overall quality across the segments went 3.01 → 3.29 (+0.28) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…4 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 4identity cos to seg 1 0.859 → 0.899 +0.039identity cos neighbours 0.933 → 0.941d_b rescored +0.291 → +0.281d_a rescored -0.266 → -0.349d_a mined -0.266d_b mined 0.292min_cos_consec (site) 0.9411min_cos_anchor (site) 0.9002dataset eurospeechlang enspeaker uk_uk_27_09122024total 60.9schain gain +2.8 dBseam step 0.8 dBcrossfades 150/150/150 ms
Script — 4 chunks, 2 with a non-speech sound
Unchanged across all 4 clips: a middle-aged masculine voice · neutral-toned, neutral-bright, fairly smooth, normally alert, slightly relaxed, fairly steady, light breath
(thankfulness gratitude, affection, interest · measured, some disfluency, average clarity, formal) Bill Dan Jarvis: My hon. Friend raises an important point. One of the most humbling parts of this job is meeting those who have been the victims of terrorism (ahem) and their families.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as thankfulness gratitude, affection, interest; style: formal, monologue; average recording, quiet background; genuineness 2.4/6; vocal-burst blend 0.0/10; 13.0s, EN.
uk_uk_27_09122024_20786960_20799951 · in -25.2 dBFS · gain +5.2 dB · eurospeech-00985
(thankfulness gratitude, hope enthusiasm optimism · measured, some disfluency, average clarity, monologue) to (ahem) progress this important work, (low mumble) and I intend to meet victims and survivors in the new year to hear more about their experiences and say more about what we will do as a Government to support them.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, full; average clarity, some disfluency, fairly narrow pitch, light breath; affect is neutral, neutral stance, slightly guarded; reads as thankfulness gratitude, hope enthusiasm optimism; style: monologue; average recording, no background noise; genuineness 2.3/6; vocal-burst blend 1.5/10; 15.4s, EN.
uk_uk_27_09122024_20810144_20825584 · in -25.2 dBFS · gain +5.2 dB · eurospeech-00985
(concentration, thankfulness gratitude · normal-paced, some disfluency, average clarity, monologue) to support them. The Bill will improve protective security and organisational preparedness across the UK, making us safer. We heard about the excellent work that many businesses and organisations already do to improve their security and
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as concentration, thankfulness gratitude; style: monologue; average recording, no background noise; genuineness 2.9/6; vocal-burst blend 2.1/10; 16.2s, EN.
uk_uk_27_09122024_20825584_20841744 · in -24.6 dBFS · gain +4.6 dB · eurospeech-00985
(concentration, anger, malevolence malice · measured, almost no disfluency, clear, formal) and preparedness. However, without a legislative requirement, there is no consistency. The Bill seeks to address that gap and complement the outstanding work that the police, the security services and other partners continue to do to combat the terror threat.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, anger, malevolence malice; style: formal, newsreading; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 0.0/10; 16.9s, EN.
uk_uk_27_09122024_20841744_20858608 · in -26.1 dBFS · gain +6.0 dB · eurospeech-00985
Doubt ↓  /  Contemptidentity +0.30 emotion REVERSED   c-eurospeech-AB2 · #18

This chain comes from the two-sided rule: it only counts if both emotions move — Doubt down and Contempt up — by at least 0.25 each.

The chain starts with Contempt around average — 0.47, lower than 53 % of clips in this corpus — and ends with it at the very top of the corpus at 0.94, higher than 94 % of clips in this corpus. That is a total rise of 0.47.

At the same time Doubt goes the other way, from 0.89 (higher than 89 % of clips in this corpus) to 0.64 (higher than 64 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.14, then +0.11, then +0.14, then +0.08 — a fairly even climb, though some clips carry more of the change than others.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.20 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.13 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.20, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 69 s · bg · eurospeech

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.202 before conversion and 0.501 after — it rose by 0.300. Neighbour-to-neighbour the worst pair went 0.117 → 0.736. (The earlier render, with segment 1 left raw, scores 0.503 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

The emotional move did not survive. Re-scored end to end, Contempt moved +0.469 in the original and -0.317 after conversion — it changed direction. On this chain the corrected audio is not an improvement. On the other named axis, Doubt, -0.258 became -0.328.

Quality. Mean predicted overall quality across the segments went 3.01 → 3.22 (+0.21) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.202 → 0.501 +0.300identity cos neighbours 0.117 → 0.736d_b rescored +0.469 → -0.317d_a rescored -0.258 → -0.328d_a mined -0.258d_b mined 0.469min_cos_consec (site) 0.1286min_cos_anchor (site) 0.1982dataset eurospeechlang bgspeaker bulgaria_bulgaria_0_151220total 68.0schain gain +2.5 dBseam step 4.8 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 3 with a non-speech sound
Unchanged across all 5 clips: a middle-aged masculine voice
(slow, very low-energy, relaxed, conversational) Заповядайте, госпожо Димова.
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, slow, relaxed, fairly steady; timbre is slightly cool, very dark, rough, thin; slurred, frequent disfluency, narrow pitch range, normal breath; affect is negative, submissive, vulnerable; no dominant emotion; style: conversational, dramatic; below-average recording, no background noise; genuineness 1.9/6; vocal-burst blend 0.0/10; 12.5s, BG.
bulgaria_bulgaria_0_15122022_10542815_10555272 · in -42.9 dBFS · gain +22.9 dB · eurospeech-00044
(pride, relief, contentment · measured, very low-energy, relaxed, monologue) Колеги, уведомявам Ви, че към 12,20 ч. (low mumble) ще дам почивка, така че Ви призовавам към експедитивна работа.
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, measured, relaxed, steady; timbre is slightly warm, slightly dark, slightly rough, slightly thin; slurred, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly negative, submissive, neutral openness; reads as pride, relief, contentment; style: monologue, whispered; below-average recording, quiet background; genuineness 2.1/6; vocal-burst blend 0.0/10; 12.3s, BG.
bulgaria_bulgaria_0_15122022_10555272_10567536 · in -37.7 dBFS · gain +17.7 dB · eurospeech-00044
(thankfulness gratitude, pride, shame · fast, normally alert, neutral tension, casual) Благодаря, господин Председател. Уважаеми колеги, аз ще бъда изключително кратка. Искам само да обоснова моето предложение по реда на Правилника, което направих в Комисията.
full caption & clip details
A child somewhat masculine voice; delivery is normally alert, fast, neutral tension, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, slightly thin; somewhat unclear, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as thankfulness gratitude, pride, shame; style: casual, dramatic; average recording, quiet background; genuineness 3.5/6; vocal-burst blend 5.0/10; 12.4s, BG.
bulgaria_bulgaria_0_15122022_10567536_10579887 · in -32.1 dBFS · gain +12.1 dB · eurospeech-00044
(disgust, disappointment, bitterness · brisk, normally alert, neutral tension, casual) (ahem) Преди това искам да благодаря на всички колеги, които подкрепихме това предложение в Комисията, свързано със ставката 9% на печатни и периодични издания, вестници и списания на физически носители, извършвана по електронен път или и двете.
full caption & clip details
A child feminine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, slightly thin; somewhat unclear, some disfluency, wide pitch range, normal breath; affect is positive, slightly dominant, slightly guarded; reads as disgust, disappointment, bitterness; style: casual, cartoonish; average recording, some background noise; genuineness 4.1/6; vocal-burst blend 7.8/10; 18.1s, BG.
bulgaria_bulgaria_0_15122022_10579887_10598032 · in -33.5 dBFS · gain +13.5 dB · eurospeech-00044
(contempt, disgust · normal-paced, normally alert, neutral tension, casual) Това е промяната, която (low mumble) предложих. Тук е мястото да кажа и на господин Сабрутев: на Комисията се разбрахме, че може да бъде и отделна точка, но референтите след това – Вие сте го разбрали явно,
full caption & clip details
A child feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, slightly thin; average clarity, some disfluency, wide pitch range, light breath; affect is positive, slightly dominant, slightly guarded; reads as contempt, disgust; style: casual, cartoonish; average recording, quiet background; genuineness 3.8/6; vocal-burst blend 6.5/10; 13.6s, BG.
bulgaria_bulgaria_0_15122022_10598032_10611599 · in -33.0 dBFS · gain +13.0 dB · eurospeech-00044
Pride ↓  /  Hope Enthusiasm Optimismidentity +0.04 emotion 35 %   c-eurospeech-AB2 · #19

This chain comes from the two-sided rule: it only counts if both emotions move — Pride down and Hope Enthusiasm Optimism up — by at least 0.25 each.

The chain starts with Hope Enthusiasm Optimism clearly present — 0.62, higher than 62 % of clips in this corpus — and ends with it at the very top of the corpus at 0.91, higher than 91 % of clips in this corpus. That is a total rise of 0.29.

At the same time Pride goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.73 (higher than 73 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.

It takes 3 clips to get there. Clip to clip the moves are +0.06, then +0.22 — a slow start, with most of the change arriving in the final step.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.86 against the first clip, where 1.00 would mean an identical voice. That is above the 0.80 threshold the mining used — very likely one person. Neighbouring clips score at worst 0.82 against each other.

Voice consistency: these clips are separate recordings joined together, matching at 0.86. You may notice the voice shift a little between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

3 clips · 44 s · hr · eurospeech

What was done to this chain. All 3 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.817 before conversion and 0.860 after — it rose by 0.043. Neighbour-to-neighbour the worst pair went 0.810 → 0.864. (The earlier render, with segment 1 left raw, scores 0.602 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Hope Enthusiasm Optimism moved +0.285 in the original and +0.099 after conversion — 35 % of the delta retained, so a meaningful part of the trajectory was flattened. On the other named axis, Pride, -0.258 became +0.003.

Quality. Mean predicted overall quality across the segments went 3.22 → 3.42 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…3 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 3identity cos to seg 1 0.817 → 0.860 +0.043identity cos neighbours 0.810 → 0.864d_b rescored +0.285 → +0.099d_a rescored -0.258 → +0.003d_a mined -0.257d_b mined 0.286min_cos_consec (site) 0.8152min_cos_anchor (site) 0.8563dataset eurospeechlang hrspeaker croatia_20140327093110-103total 42.8schain gain +2.7 dBseam step 1.3 dBcrossfades 150/150 ms
Script — 3 chunks, 1 with a non-speech sound
Unchanged across all 3 clips: an adult masculine voice · neutral-bright, slightly rough, balanced body, average recording, quiet background, neutral tension, wide pitch range, normal breath
(pride, anger, impatience and irritability · normal-paced, energised, moderately variable, casual) Slažem se s vama, ali morate znati da je kapital anacionalan i kad sam rekao da su oni štitili samo najveće, štitili su one koji imaju najveće novce, najviše kapitala, naravno da su štitili
full caption & clip details
An adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, fairly guarded; reads as pride, anger, impatience and irritability; style: casual, storytelling; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 5.7/10; 11.6s, HR.
croatia_20140327093110-10313_3152976_3164560 · in -21.1 dBFS · gain +1.1 dB · eurospeech-01385
(contempt, anger, concentration · normal-paced, energised, moderately variable, authoritative) kapital, a kapital je anacionalan i onda je agencija, slažem se s vama u potpunosti, anacionalna jer ona nije štitila tržište Hrvatske, ona nije štitila interese jer to su znači štititi interese Republike Hrvatske.
full caption & clip details
A middle-aged masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, wide pitch range, normal breath; affect is neutral, slightly dominant, fairly guarded; reads as contempt, anger, concentration; style: authoritative, monologue; average recording, quiet background; genuineness 3.3/6; vocal-burst blend 4.2/10; 16.3s, HR.
croatia_20140327093110-10313_3164560_3180848 · in -21.2 dBFS · gain +1.2 dB · eurospeech-01385
(hope enthusiasm optimism · measured, subdued, fairly steady, monologue) (low mumble) Ovo što ste govorili o povlačenju novca iz EU fondova, to se dvije godine već govori. Nećete povući, ne radite dobro, ali šta je problem, u softveru. Pazite, kažu da softver ne radi.
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, neutral tension, fairly steady; timbre is slightly cool, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, wide pitch range, normal breath; affect is mildly negative, slightly dominant, fairly guarded; reads as hope enthusiasm optimism; style: monologue, casual; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 2.1/10; 15.3s, HR.
croatia_20140327093110-10313_3180848_3196192 · in -20.8 dBFS · gain +0.8 dB · eurospeech-01385
Confusion ↓  /  Concentrationidentity +0.37 emotion 161 %   c-eurospeech-AB2 · #20

This chain comes from the two-sided rule: it only counts if both emotions move — Confusion down and Concentration up — by at least 0.25 each.

The chain starts with Concentration clearly present — 0.67, higher than 67 % of clips in this corpus — and ends with it at the very top of the corpus at 0.97, higher than 97 % of clips in this corpus. That is a total rise of 0.31.

At the same time Confusion goes the other way, from 0.98 (higher than 98 % of clips in this corpus) to 0.71 (higher than 71 % of clips in this corpus), a change of -0.26. Both halves had to happen for this chain to qualify.

It takes 5 clips to get there. Clip to clip the moves are +0.14, then +0.18, then -0.15, then +0.14 — not a clean run: step 3 moves back the other way by 0.15 before the chain recovers.

No single step is larger than the 0.25 cap, which is exactly what stops this being a jump cut: the change has to be spread across the clips instead of landing all at once.

Same speaker? The least similar clip scores 0.45 against the first clip, where 1.00 would mean an identical voice. That is below the 0.80 threshold the mining used — treat the “same speaker” claim here with caution. Neighbouring clips score at worst 0.38 against each other.

Voice consistency: these clips are separate recordings joined together and the match is loose (0.45, under the 0.80 threshold), so the voice may audibly change between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

5 clips · 69 s · da · eurospeech

What was done to this chain. All 5 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.430 before conversion and 0.800 after — it rose by 0.370. Neighbour-to-neighbour the worst pair went 0.376 → 0.885. (The earlier render, with segment 1 left raw, scores 0.627 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Concentration moved +0.307 in the original and +0.494 after conversion — 161 % of the delta retained, i.e. the move came out slightly larger after conversion than before. On the other named axis, Confusion, -0.263 became -0.226.

Quality. Mean predicted overall quality across the segments went 3.23 → 3.46 (+0.22) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…5 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 5identity cos to seg 1 0.430 → 0.800 +0.370identity cos neighbours 0.376 → 0.885d_b rescored +0.307 → +0.494d_a rescored -0.263 → -0.226d_a mined -0.264d_b mined 0.307min_cos_consec (site) 0.3815min_cos_anchor (site) 0.4472dataset eurospeechlang daspeaker denmark_20222M058_2023-05-total 67.5schain gain +3.8 dBseam step 1.0 dBcrossfades 150/150/150/150 ms
Script — 5 chunks, 1 with a non-speech sound
Unchanged across all 5 clips: a young adult masculine voice · fairly smooth, balanced body, average recording, normally alert
(confusion, intoxication altered states of consciousness, shame · slow, slightly relaxed, fairly steady, didactic) Støjberg. Kl. 14:31 Inger Støjberg (DD): Det er jeg ikke enig i. For hvis man antager, at ministeren har ret,
full caption & clip details
A young adult masculine voice; delivery is normally alert, slow, slightly relaxed, fairly steady; timbre is slightly cool, slightly dark, fairly smooth, balanced body; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as confusion, intoxication altered states of consciousness, shame; style: didactic, whispered; average recording, quiet background; genuineness 3.6/6; vocal-burst blend 1.1/10; 10.4s, DA.
denmark_20222M058_2023-05-10_1000_16352880_16363328 · in -23.0 dBFS · gain +3.0 dB · eurospeech-00482
(thankfulness gratitude, relief · measured, slightly relaxed, fairly steady, monologue) så siger ministeren jo også, at det sammenhold, som et kæmpestort flertal i Danmark har (ahem) om, at vi netop har ytringsfrihed, at vi har religionsfrihed, men ikke religionslighed, og at vi har også ret til at kritisere religion,
full caption & clip details
A young adult feminine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as thankfulness gratitude, relief; style: monologue, whispered; average recording, no background noise; genuineness 2.6/6; vocal-burst blend 5.3/10; 11.3s, DA.
denmark_20222M058_2023-05-10_1000_16363328_16374655 · in -22.3 dBFS · gain +2.3 dB · eurospeech-00482
(bitterness, concentration, jealousy and envy · normal-paced, slightly relaxed, fairly steady, monologue) at kritisere religion, ikke har en betydning. Selvfølgelig har det en betydning, selvfølgelig er det et rygstød, som hr. Morten Messerschmidt sagde. Selvfølgelig er det et rygstød, som også hr. Jacob Mark sagde, til de lærere, der står og skal undervise i det her, at man rent faktisk fra Folketingets side, fra samfundets side
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as bitterness, concentration, jealousy and envy; style: monologue, whispered; average recording, quiet background; genuineness 2.8/6; vocal-burst blend 4.3/10; 17.1s, DA.
denmark_20222M058_2023-05-10_1000_16374655_16391743 · in -20.5 dBFS · gain +0.5 dB · eurospeech-00482
(pride, triumph, confusion · brisk, neutral tension, moderately variable, conversational) har en opbakning til, at det her selvfølgelig også skal være en del af undervisningen. Hvis ikke man tror på sammenholdet, er det fint nok, hvad ministeren siger, men hvis man tror på sammenholdet – det gør jeg,
full caption & clip details
An adult masculine voice; delivery is normally alert, brisk, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as pride, triumph, confusion; style: conversational, casual; average recording, quiet background; genuineness 3.2/6; vocal-burst blend 7.4/10; 11.2s, DA.
denmark_20222M058_2023-05-10_1000_16391743_16402991 · in -21.5 dBFS · gain +1.5 dB · eurospeech-00482
(concentration, triumph, bitterness · normal-paced, slightly relaxed, fairly steady, monologue) om, at lærere skal kunne udøve deres lærergerning på en måde, hvor de får videreleveret hele vores kulturhistorie, inklusive den nyere danmarkshistorie, står jeg selvfølgelig et hundrede ti procent bag. Det, jeg også bare siger, er, at hvis opgaven er at nå derhen, skal vi jo interessere os for, om den politik, vi vedtager, så bringer os derhen eller ej.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as concentration, triumph, bitterness; style: monologue, didactic; average recording, quiet background; genuineness 1.6/6; vocal-burst blend 5.1/10; 18.2s, DA.
denmark_20222M058_2023-05-10_1000_16416496_16434655 · in -23.6 dBFS · gain +3.5 dB · eurospeech-00482