rescue rule sad-Awe-S3-k2 — voice-corrected

Awe under rescue rule S3, k=2. does the winning rule generalise? Awe has the widest gap (0.474). 95.1 % of clips score at or below zero on this emotion and the largest gap on its normalised axis is 0.474 (WIDER than the 0.25 step cap). This rule found 5,762 chains over 40,000 tracks; the strict rule found 0 at k=3.

This is the voice-conversion-corrected version of this tier. The original, uncorrected page is still there and unchanged: t_sad-Awe-S3-k2.html. Every card below carries both renders so you can switch between them without leaving the page. What was done, and what it measured →
These are not strict-rule trajectories. They come from a deliberately looser rule, built to recover examples on an emotion the strict rule cannot reach. What rule S3 changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'. What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained. Full explanation →
20chains converted
20segments re-voiced
0.795 → 0.790median worst-to-anchor identity cosine
99 %median emotion delta retained
How to read a Script. Each chunk is one line: a short tag of what the models heard in that clip, then the words spoken.

(underlined, plain · delivery, style) — the tag before the words. Emotions first, then how it is delivered. Underlined descriptors are the ones that change across this chain — anything identical on every clip is pulled out and stated once above.
(ahem) — brackets inside the words are a real non-speech sound, printed where it happens.

The full generated caption for any clip is under “full caption & clip details”. Its perceived-gender and background-noise clauses were re-rendered from the numeric buckets, because the versions stored in the corpus index had those two ladders running backwards.

The emotion clause has been re-derived, and it used to be wrong. The 40 emotion heads are not on a common scale — Interest has a median of 2.08 and is never zero, while Infatuation is zero on 87.7 % of clips — and the caption named an emotion whenever its raw score cleared an absolute 1.0. Interest therefore appeared in 94.8 % of captions and Sadness in almost none: the clause was reporting the scale of the head, not the emotion of the clip. An emotion is now named only when it lands in the top 10 % for that emotion, against a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning every dataset and language. Interest now appears in 6.2 %, all 40 emotions occur, and a clip that is ordinary on all 40 says “no dominant emotion” rather than being forced to pick one (21.8 % of clips). This is the same scale the trajectory miner selects on, so the caption and the mining now refer to the same quantity: the mined target emotion is named in the final clip's caption on 73 % of chains, up from 46 %.

The tags and captions describe the ORIGINAL clips — they are the corpus annotation, not a re-reading of the converted audio. The re-scored numbers for the converted audio are the ones printed in each card's own paragraph and chips.
Awe rising ↑identity −0.01 emotion 98 %   sad-Awe-S3-k2 · #1

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Awe.

The raw scorer output across the chain is -0.00, 1.31. In the first clip the scorer found no Awe whatsoever (-0.00); by the last it is at 1.31.

On the corpus-wide percentile scale those become 0.47, 1.00 — a total move of +0.52.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Awe, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Awe sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 27 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.923 before conversion and 0.909 after — it fell by 0.014. Neighbour-to-neighbour the worst pair went 0.923 → 0.909. (The earlier render, with segment 1 left raw, scores 0.818 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Awe moved +0.523 in the original and +0.511 after conversion — 98 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 3.16 → 3.28 (+0.12) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores -0.0 → 1.3076normalised 0.477 → 0.997identity cos to seg 1 0.923 → 0.909 -0.014identity cos neighbours 0.923 → 0.909d_b rescored +0.523 → +0.511d_a rescored +0.523 → +0.511d_a mined 1.308d_b mined 0.520min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_163DYBR_hRgtotal 26.3schain gain +2.5 dBseam step 1.0 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(steady, formal, monologue) Startschuss für DSDS. Nicht mehr lange, und Florian Silbereisen-Fans kommen auch bei RTL auf ihre Kosten
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; no dominant emotion; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 7.9s, DE.
DE_163DYBR_hRg_W000000 · in -16.4 dBFS · gain -3.6 dB · emolia-00161
(awe, pain, fatigue exhaustion · fairly steady, narration, formal) Ich bin für alles offen und am Ende ist es egal, ob ein Kandidat Schlager, Pop oder Rock singt. Hauptsache, derjenige kann gut singen. Denn wir wollen ja am Ende einen Superstar finden, die Musikrichtung ist dabei egal. Das Unterhosengeheimnis von Florian Silbereisen.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as awe, pain, fatigue exhaustion; style: narration, formal; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 18.6s, DE.
DE_163DYBR_hRg_W000010 · in -15.9 dBFS · gain -4.1 dB · emolia-00161
Awe rising ↑identity −0.04 emotion 100 %   sad-Awe-S3-k2 · #2

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Awe.

The raw scorer output across the chain is -0.00, 1.16. In the first clip the scorer found no Awe whatsoever (-0.00); by the last it is at 1.16.

On the corpus-wide percentile scale those become 0.47, 1.00 — a total move of +0.52.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Awe, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Awe sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 17 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.848 before conversion and 0.805 after — it fell by 0.043. Neighbour-to-neighbour the worst pair went 0.848 → 0.805. (The earlier render, with segment 1 left raw, scores 0.718 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Awe moved +0.520 in the original and +0.521 after conversion — 100 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.89 → 3.14 (+0.25) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores -0.0 → 1.1631normalised 0.477 → 0.995identity cos to seg 1 0.848 → 0.805 -0.043identity cos neighbours 0.848 → 0.805d_b rescored +0.520 → +0.521d_a rescored +0.520 → +0.521d_a mined 1.163d_b mined 0.518min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_1A7Boronwiktotal 16.6schain gain +1.3 dBseam step 0.2 dBcrossfades 100 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an elderly somewhat feminine voice · slightly warm, smooth, balanced body, good recording, no background noise, slow, very low-energy, relaxed
(contemplation, pain, emotional numbness · frequent disfluency, fairly narrow pitch, ASMR, monologue) Ein Gebet, was das Gleichgewicht zwischen geben und nehmen herstellt. Das wichtige Gleichgewicht.
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, slow, relaxed, steady; timbre is slightly warm, slightly dark, smooth, balanced body; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly positive, submissive, slightly vulnerable; reads as contemplation, pain, emotional numbness; style: ASMR, monologue; good recording, no background noise; genuineness 2.5/6; vocal-burst blend 2.5/10; 8.5s, DE.
DE_1A7Boronwik_W000002 · in -17.6 dBFS · gain -2.4 dB · emolia-00254
(awe, relief, contentment · little disfluency, narrow pitch range, ASMR, whispered) Dass wir mithilfe Gottes ganz viel Frieden und Licht auf die Welt bringen werden.
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, slow, relaxed, steady; timbre is slightly warm, dark, smooth, balanced body; somewhat unclear, little disfluency, narrow pitch range, audible breath; affect is mildly positive, submissive, neutral openness; reads as awe, relief, contentment; style: ASMR, whispered; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 3.3/10; 8.3s, DE.
DE_1A7Boronwik_W000026 · in -18.1 dBFS · gain -1.9 dB · emolia-00254
Awe rising ↑identity +0.60 emotion 99 %   sad-Awe-S3-k2 · #3

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Awe.

The raw scorer output across the chain is -0.00, 1.10. In the first clip the scorer found no Awe whatsoever (-0.00); by the last it is at 1.10.

On the corpus-wide percentile scale those become 0.47, 0.99 — a total move of +0.52.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Awe, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Awe sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 10 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was -0.058 before conversion and 0.540 after — it rose by 0.599. Neighbour-to-neighbour the worst pair went -0.058 → 0.540. (The earlier render, with segment 1 left raw, scores 0.168 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Awe moved +0.519 in the original and +0.515 after conversion — 99 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.65 → 2.80 (+0.15) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores -0.0 → 1.0957normalised 0.477 → 0.994identity cos to seg 1 -0.058 → 0.540 +0.599identity cos neighbours -0.058 → 0.540d_b rescored +0.519 → +0.515d_a rescored +0.519 → +0.515d_a mined 1.096d_b mined 0.517min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_1Pd71w8hAhItotal 9.6schain gain +2.7 dBseam step 2.0 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a young adult feminine voice · neutral-toned, neutral-bright, fairly smooth, quiet background, normal-paced, normally alert, some disfluency, moderate pitch range
(embarrassment, fatigue exhaustion, infatuation · neutral tension, moderately variable, somewhat unclear, casual) Ich drück mich immer falsch auf, aus, in dieser, also, filterraum, ich muss auch mal die richtige Sprache lernen.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; somewhat unclear, some disfluency, moderate pitch range, normal breath; affect is positive, slightly submissive, slightly vulnerable; reads as embarrassment, fatigue exhaustion, infatuation; style: casual, conversational; below-average recording, quiet background; genuineness 5.3/6; vocal-burst blend 4.9/10; 6.3s, DE.
DE_1Pd71w8hAhI_W000000 · in -21.4 dBFS · gain +1.4 dB · emolia-00077
(astonishment surprise, awe, longing · slightly relaxed, fairly steady, average clarity, casual) Und dann noch diese Erhöhungen, die da immer kommen, mein Gott, ja.
full caption & clip details
A young adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly positive, slightly submissive, neutral openness; reads as astonishment surprise, awe, longing; style: casual, conversational; average recording, quiet background; genuineness 4.8/6; vocal-burst blend 3.2/10; 3.5s, DE.
DE_1Pd71w8hAhI_W000187 · in -19.4 dBFS · gain -0.6 dB · emolia-00077
Awe rising ↑identity −0.10 emotion REVERSED   sad-Awe-S3-k2 · #4

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Awe.

The raw scorer output across the chain is -0.00, 1.32. In the first clip the scorer found no Awe whatsoever (-0.00); by the last it is at 1.32.

On the corpus-wide percentile scale those become 0.47, 1.00 — a total move of +0.52.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Awe, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Awe sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 20 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.822 before conversion and 0.725 after — it fell by 0.097. Neighbour-to-neighbour the worst pair went 0.822 → 0.725. (The earlier render, with segment 1 left raw, scores 0.597 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Awe moved +0.523 in the original and +0.000 after conversion — 0 % of the delta retained, so a meaningful part of the trajectory was flattened.

Quality. Mean predicted overall quality across the segments went 3.08 → 3.22 (+0.14) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores -0.0 → 1.3232normalised 0.477 → 0.998identity cos to seg 1 0.822 → 0.725 -0.097identity cos neighbours 0.822 → 0.725d_b rescored +0.523 → +0.000d_a rescored +0.523 → +0.000d_a mined 1.323d_b mined 0.520min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_1aPh9BCqtRutotal 19.2schain gain +0.8 dBseam step 2.4 dBcrossfades 100 ms
Script — 2 chunks, 2 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · neutral-bright, fairly smooth, balanced body, average recording, quiet background, normal-paced, energised, moderately variable
(disgust, interest, teasing · neutral tension, frequent disfluency, very clear, storytelling) 3G Symmension 3. Ihr wisst Bescheid. Ganz genau. Und heute befinden wir uns, wenn ich das hier alles richtig sehe, und mich gut informiert hab, in Pyra Media. Ja, (ahem) ganz genau.
full caption & clip details
A young adult masculine voice; delivery is energised, normal-paced, neutral tension, moderately variable; timbre is slightly cool, neutral-bright, fairly smooth, balanced body; very clear, frequent disfluency, very wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as disgust, interest, teasing; style: storytelling, playful; average recording, quiet background; genuineness 3.7/6; vocal-burst blend 1.9/10; 14.1s, DE.
DE_1aPh9BCqtRu_W000000 · in -15.4 dBFS · gain -4.6 dB · emolia-00022
(astonishment surprise, awe, confusion · slightly relaxed, some disfluency, average clarity, cartoonish) Oh mein Gott! Ja, (ahem) macht ja auch Sinn, so dieses Pyramidenzeugs.
full caption & clip details
A child masculine voice; delivery is energised, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, very wide pitch range, normal breath; affect is positive, slightly dominant, neutral openness; reads as astonishment surprise, awe, confusion; style: cartoonish, dramatic; average recording, quiet background; genuineness 3.4/6; vocal-burst blend 0.4/10; 5.3s, DE.
DE_1aPh9BCqtRu_W000001 · in -15.5 dBFS · gain -4.5 dB · emolia-00022
Awe rising ↑identity +0.48 emotion 100 %   sad-Awe-S3-k2 · #5

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Awe.

The raw scorer output across the chain is -0.00, 1.16. In the first clip the scorer found no Awe whatsoever (-0.00); by the last it is at 1.16.

On the corpus-wide percentile scale those become 0.47, 1.00 — a total move of +0.52.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Awe, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Awe sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 10 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.262 before conversion and 0.747 after — it rose by 0.484. Neighbour-to-neighbour the worst pair went 0.262 → 0.747. (The earlier render, with segment 1 left raw, scores 0.682 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Awe moved +0.520 in the original and +0.519 after conversion — 100 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.79 → 3.01 (+0.22) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores -0.0 → 1.1611normalised 0.477 → 0.995identity cos to seg 1 0.262 → 0.747 +0.484identity cos neighbours 0.262 → 0.747d_b rescored +0.520 → +0.519d_a rescored +0.520 → +0.519d_a mined 1.161d_b mined 0.518min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_1eb_VRL8MFItotal 9.5schain gain +1.5 dBseam step 0.8 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, good recording, no background noise, normal-paced, normally alert
(affection · wide pitch range, authoritative, storytelling) Tage der Ermutigung. Ich freue mich, dass du wieder eingeschaltet hast, dass du wieder dabei bist.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, wide pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as affection; style: authoritative, storytelling; good recording, no background noise; genuineness 1.4/6; vocal-burst blend 0.0/10; 4.9s, DE.
DE_1eb_VRL8MFI_W000000 · in -19.3 dBFS · gain -0.7 dB · emolia-00216
(awe, elation, pleasure ecstasy · moderate pitch range, formal, storytelling) Das war der Anfang einer wunderbaren Reise mit Jesus Christus.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as awe, elation, pleasure ecstasy; style: formal, storytelling; good recording, no background noise; genuineness 0.7/6; vocal-burst blend 0.0/10; 4.9s, DE.
DE_1eb_VRL8MFI_W000268 · in -18.9 dBFS · gain -1.1 dB · emolia-00216
Awe rising ↑identity −0.05 emotion 98 %   sad-Awe-S3-k2 · #6

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Awe.

The raw scorer output across the chain is -0.00, 1.42. In the first clip the scorer found no Awe whatsoever (-0.00); by the last it is at 1.42.

On the corpus-wide percentile scale those become 0.47, 1.00 — a total move of +0.52.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Awe, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Awe sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 20 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.724 before conversion and 0.675 after — it fell by 0.049. Neighbour-to-neighbour the worst pair went 0.724 → 0.675. (The earlier render, with segment 1 left raw, scores 0.605 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Awe moved +0.524 in the original and +0.512 after conversion — 98 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.97 → 3.23 (+0.25) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores -0.0 → 1.418normalised 0.477 → 0.998identity cos to seg 1 0.724 → 0.675 -0.049identity cos neighbours 0.724 → 0.675d_b rescored +0.524 → +0.512d_a rescored +0.524 → +0.512d_a mined 1.418d_b mined 0.521min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_2UhFJjC9wtMtotal 20.1schain gain -0.3 dBseam step 1.7 dBcrossfades 150 ms
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, average recording, quiet background, normally alert, moderately variable
(intoxication altered states of consciousness, elation, hope enthusiasm optimism · normal-paced, neutral tension, average clarity, casual) Hallo und willkommen hier zu F1 2019. 哈喽,欢迎来到F1 2019。嗯, (low mumble) ja, ich werde keine Karriere heute starten, sondern eine Formel 2 Meisterschaft, weil ich habe gemerkt, die Fahrzeuge sind mega mega nice und, 嗯, (low mumble)
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, wide pitch range, light breath; affect is positive, slightly submissive, slightly guarded; reads as intoxication altered states of consciousness, elation, hope enthusiasm optimism; style: casual, playful; average recording, quiet background; genuineness 4.2/6; vocal-burst blend 1.8/10; 15.7s, DE.
DE_2UhFJjC9wtM_W000000 · in -14.5 dBFS · gain -5.5 dB · emolia-00216
(awe, astonishment surprise, intoxication altered states of consciousness · measured, slightly relaxed, somewhat unclear, casual) Wow, Aitken hat da extrem früh gebremst und Alvin führt das Rennen an.
full caption & clip details
A young adult masculine voice; delivery is normally alert, measured, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as awe, astonishment surprise, intoxication altered states of consciousness; style: casual, conversational; average recording, quiet background; genuineness 4.5/6; vocal-burst blend 0.0/10; 4.7s, DE.
DE_2UhFJjC9wtM_W000036 · in -14.8 dBFS · gain -5.2 dB · emolia-00216
Awe rising ↑identity +0.62 emotion 98 %   sad-Awe-S3-k2 · #7

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Awe.

The raw scorer output across the chain is -0.00, 1.19. In the first clip the scorer found no Awe whatsoever (-0.00); by the last it is at 1.19.

On the corpus-wide percentile scale those become 0.47, 1.00 — a total move of +0.52.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Awe, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Awe sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 28 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.068 before conversion and 0.693 after — it rose by 0.624. Neighbour-to-neighbour the worst pair went 0.068 → 0.693. (The earlier render, with segment 1 left raw, scores 0.703 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Awe moved +0.521 in the original and +0.508 after conversion — 98 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.86 → 3.00 (+0.14) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores -0.0 → 1.1865normalised 0.477 → 0.996identity cos to seg 1 0.068 → 0.693 +0.624identity cos neighbours 0.068 → 0.693d_b rescored +0.521 → +0.508d_a rescored +0.521 → +0.508d_a mined 1.187d_b mined 0.519min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_2_pDO2tImiMtotal 27.7schain gain +4.1 dBseam step 1.6 dBcrossfades 100 ms
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: a child feminine voice · neutral-toned, neutral-bright, balanced body, average recording, some background noise, light breath
(pride, affection, elation · normal-paced, normally alert, slightly relaxed, casual) (ahem) Karl gehört, (ahem) ähm, meinem Mann und mir. Der wohnt, (low mumble) ähm, mit uns auf einem kleinen Bauernhof in, (low mumble) ähm, Bitburg-Hau. Das ist unser Pit.
full caption & clip details
A child feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, frequent disfluency, wide pitch range, light breath; affect is positive, slightly submissive, neutral openness; reads as pride, affection, elation; style: casual, conversational; average recording, some background noise; genuineness 4.1/6; vocal-burst blend 1.2/10; 9.4s, DE.
DE_2_pDO2tImiM_W000000 · in -16.4 dBFS · gain -3.5 dB · emolia-00196
(awe, pride, sadness · measured, subdued, neutral tension, monologue) Wo er alle Dinge, die Gott geschaffen hat, also die Sonne, den Mond, die Sterne, das Wasser, die Berge, die Blumen, die Tiere, die Menschen, alle seine Brüder und Schwestern nennt. Alle sind Geschöpfe Gottes. Gott ist unser Allervater. Von daher haben wir eine ganz besondere Affinität auch zur Tiersegel.
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as awe, pride, sadness; style: monologue; average recording, some background noise; genuineness 3.5/6; vocal-burst blend 2.2/10; 18.5s, DE.
DE_2_pDO2tImiM_W000029 · in -21.4 dBFS · gain +1.4 dB · emolia-00196
Awe rising ↑identity +0.56 emotion 99 %   sad-Awe-S3-k2 · #8

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Awe.

The raw scorer output across the chain is -0.00, 1.12. In the first clip the scorer found no Awe whatsoever (-0.00); by the last it is at 1.12.

On the corpus-wide percentile scale those become 0.47, 0.99 — a total move of +0.52.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Awe, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Awe sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 20 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.227 before conversion and 0.787 after — it rose by 0.559. Neighbour-to-neighbour the worst pair went 0.227 → 0.787. (The earlier render, with segment 1 left raw, scores 0.773 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Awe moved +0.519 in the original and +0.515 after conversion — 99 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.94 → 3.06 (+0.12) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores -0.0 → 1.1201normalised 0.477 → 0.995identity cos to seg 1 0.227 → 0.787 +0.559identity cos neighbours 0.227 → 0.787d_b rescored +0.519 → +0.515d_a rescored +0.519 → +0.515d_a mined 1.120d_b mined 0.517min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_2qbSrNDq5eItotal 19.5schain gain +1.2 dBseam step 0.9 dBcrossfades 100 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult feminine voice · neutral-toned, neutral-bright, balanced body, moderately variable, wide pitch range
(fear, distress, sadness · normal-paced, normally alert, slightly relaxed, formal) Viele Menschen haben Angst und wir dürfen immer wieder zu Jesus kommen und er befreit uns davon und er hilft uns.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, wide pitch range, light breath; affect is mildly positive, neutral stance, neutral openness; reads as fear, distress, sadness; style: formal, monologue; good recording, no background noise; genuineness 0.4/6; vocal-burst blend 1.2/10; 7.1s, DE.
DE_2qbSrNDq5eI_W000001 · in -11.7 dBFS · gain -8.3 dB · emolia-00161
(relief, affection, thankfulness gratitude · slow, very low-energy, relaxed, monologue) Danke Herr, dass du jetzt gerade in diesem Moment diesem Mann so nahe bist. Und Herr, ich danke dir dafür. Herr, den ganzen Unterleib, Herr, was da alles durcheinander ist, was entzündet ist. Herr,
full caption & clip details
An elderly masculine voice; delivery is very low-energy, slow, relaxed, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; somewhat unclear, frequent disfluency, wide pitch range, audible breath; affect is mildly positive, slightly submissive, neutral openness; reads as relief, affection, thankfulness gratitude; style: monologue, casual; average recording, quiet background; genuineness 2.9/6; vocal-burst blend 1.6/10; 12.5s, DE.
DE_2qbSrNDq5eI_W000021 · in -17.7 dBFS · gain -2.3 dB · emolia-00161
Awe rising ↑identity +0.06 emotion 100 %   sad-Awe-S3-k2 · #9

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Awe.

The raw scorer output across the chain is -0.00, 1.02. In the first clip the scorer found no Awe whatsoever (-0.00); by the last it is at 1.02.

On the corpus-wide percentile scale those become 0.47, 0.99 — a total move of +0.52.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Awe, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Awe sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 18 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.681 before conversion and 0.742 after — it rose by 0.061. Neighbour-to-neighbour the worst pair went 0.681 → 0.742. (The earlier render, with segment 1 left raw, scores 0.553 here.) This chain started 0.50-0.70 — audibly different, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Awe moved +0.517 in the original and +0.516 after conversion — 100 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.78 → 3.12 (+0.35) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores -0.0 → 1.0205normalised 0.477 → 0.992identity cos to seg 1 0.681 → 0.742 +0.061identity cos neighbours 0.681 → 0.742d_b rescored +0.517 → +0.516d_a rescored +0.517 → +0.516d_a mined 1.020d_b mined 0.515min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_3QIMrq-rDcgtotal 18.1schain gain +3.8 dBseam step 1.3 dBcrossfades 100 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, quiet background, moderately variable, moderate pitch range
(embarrassment, shame · measured, very low-energy, neutral tension, casual) Die Sprache der Tiere, tja, neulich ständig nehme ich ja am Fenster und hab mit einer guten Freundin telefoniert, und da kann man so auf das Thema Krafttiere zu sprechen, ne?
full caption & clip details
An adult masculine voice; delivery is very low-energy, measured, neutral tension, moderately variable; timbre is neutral-toned, neutral-bright, fairly smooth, slightly thin; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is mildly negative, neutral stance, neutral openness; reads as embarrassment, shame; style: casual, monologue; average recording, quiet background; mildly explicit content; genuineness 4.7/6; vocal-burst blend 1.4/10; 13.7s, DE.
DE_3QIMrq-rDcg_W000000 · in -22.7 dBFS · gain +2.7 dB · emolia-00079
(awe, longing, infatuation · normal-paced, normally alert, slightly relaxed, conversational) Ja, und die Sekundärwicklung, ne, die hat mich nämlich auch etwas fasziniert.
full caption & clip details
A young adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, moderately variable; timbre is neutral-toned, neutral-bright, slightly rough, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is mildly negative, neutral stance, neutral openness; reads as awe, longing, infatuation; style: conversational, storytelling; good recording, quiet background; genuineness 2.4/6; vocal-burst blend 0.0/10; 4.6s, DE.
DE_3QIMrq-rDcg_W000053 · in -19.5 dBFS · gain -0.5 dB · emolia-00079
Awe rising ↑identity +0.12 emotion 3 %   sad-Awe-S3-k2 · #10

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Awe.

The raw scorer output across the chain is -0.00, 1.02. In the first clip the scorer found no Awe whatsoever (-0.00); by the last it is at 1.02.

On the corpus-wide percentile scale those become 0.47, 0.99 — a total move of +0.52.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Awe, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Awe sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 20 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.725 before conversion and 0.847 after — it rose by 0.122. Neighbour-to-neighbour the worst pair went 0.725 → 0.847. (The earlier render, with segment 1 left raw, scores 0.657 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Awe moved +0.517 in the original and +0.017 after conversion — 3 % of the delta retained, so a meaningful part of the trajectory was flattened.

Quality. Mean predicted overall quality across the segments went 2.97 → 3.21 (+0.24) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores -0.0 → 1.0225normalised 0.477 → 0.992identity cos to seg 1 0.725 → 0.847 +0.122identity cos neighbours 0.725 → 0.847d_b rescored +0.517 → +0.017d_a rescored +0.517 → +0.017d_a mined 1.022d_b mined 0.515min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_4N4Wcn9CHFototal 20.0schain gain +1.5 dBseam step 1.3 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a middle-aged masculine voice · neutral-bright, balanced body, good recording, no background noise, measured, fairly steady, clear, light breath
(contemplation, sadness, affection · subdued, relaxed, little disfluency, monologue) Es ist schwer zu erklären, was ich in Rumänien alles erleben dürfte. Natur, Herausforderungen, Leben, Gastfreundschaft, Liebe und so vieles vieles mehr. Meine Reisen bereichern meine Seele.
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, relaxed, fairly steady; timbre is slightly warm, neutral-bright, slightly rough, balanced body; clear, little disfluency, fairly narrow pitch, light breath; affect is mildly positive, neutral stance, fairly guarded; reads as contemplation, sadness, affection; style: monologue, narration; good recording, no background noise; genuineness 1.0/6; vocal-burst blend 0.9/10; 13.7s, DE.
DE_4N4Wcn9CHFo_W000000 · in -21.1 dBFS · gain +1.1 dB · emolia-00123
(longing, contemplation, sadness · normally alert, slightly relaxed, some disfluency, storytelling) Es wird immer einsamer, und es ist wie ein Traum, durch Wasserdurchquerungen mitten in der Natur zu fahren.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, some disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as longing, contemplation, sadness; style: storytelling, monologue; good recording, no background noise; genuineness 1.5/6; vocal-burst blend 0.7/10; 6.6s, DE.
DE_4N4Wcn9CHFo_W000014 · in -19.2 dBFS · gain -0.8 dB · emolia-00123
Awe rising ↑identity +0.07 emotion 100 %   sad-Awe-S3-k2 · #11

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Awe.

The raw scorer output across the chain is -0.00, 1.07. In the first clip the scorer found no Awe whatsoever (-0.00); by the last it is at 1.07.

On the corpus-wide percentile scale those become 0.47, 0.99 — a total move of +0.52.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Awe, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Awe sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 36 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.797 before conversion and 0.865 after — it rose by 0.069. Neighbour-to-neighbour the worst pair went 0.797 → 0.865. (The earlier render, with segment 1 left raw, scores 0.737 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Awe moved +0.518 in the original and +0.516 after conversion — 100 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 3.13 → 3.33 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores -0.0 → 1.0674normalised 0.477 → 0.993identity cos to seg 1 0.797 → 0.865 +0.069identity cos neighbours 0.797 → 0.865d_b rescored +0.518 → +0.516d_a rescored +0.518 → +0.516d_a mined 1.067d_b mined 0.516min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_4bFqkygOSLktotal 35.9schain gain +1.4 dBseam step 1.3 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · balanced body, no background noise
(fear, bitterness, sourness · normal-paced, normally alert, slightly relaxed, monologue) Obwohl die allgemeine Angst wuchs, dass man es mit einem Krankheitserreger zu tun hatte, der wie die Tollwut aggressiv machte und im schlimmsten Fall auch auf Menschen übertragbar war, wollte der Kollege in Wilhelmsburg nichts davon wissen.
full caption & clip details
An adult masculine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as fear, bitterness, sourness; style: monologue, narration; good recording, no background noise; genuineness 2.3/6; vocal-burst blend 0.7/10; 14.0s, DE.
DE_4bFqkygOSLk_W000000 · in -18.7 dBFS · gain -1.3 dB · emolia-00062
(awe, bitterness, contemplation · slow, very low-energy, relaxed, narration) Im Zwielicht schien sich das Zimmer, um ihn herum zu bewegen. Aber als er das Licht anknipste, brachte die Beleuchtung der Nachtiglampe die vertraute Ordnung zurück. Schrank, Stuhl, Bücherregal, Nachtig. Alles an seinem Platz. Im Bett begann er so sehr zu schwitzen,
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, slow, relaxed, steady; timbre is slightly warm, slightly dark, rough, balanced body; slurred, frequent disfluency, fairly narrow pitch, normal breath; affect is mildly negative, neutral stance, fairly guarded; reads as awe, bitterness, contemplation; style: narration, whispered; average recording, no background noise; genuineness 1.4/6; vocal-burst blend 1.7/10; 22.1s, DE.
DE_4bFqkygOSLk_W000044 · in -23.1 dBFS · gain +3.1 dB · emolia-00062
Awe rising ↑identity −0.06 emotion 100 %   sad-Awe-S3-k2 · #12

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Awe.

The raw scorer output across the chain is -0.00, 1.16. In the first clip the scorer found no Awe whatsoever (-0.00); by the last it is at 1.16.

On the corpus-wide percentile scale those become 0.47, 1.00 — a total move of +0.52.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Awe, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Awe sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 20 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.831 before conversion and 0.774 after — it fell by 0.057. Neighbour-to-neighbour the worst pair went 0.831 → 0.774. (The earlier render, with segment 1 left raw, scores 0.706 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Awe moved +0.520 in the original and +0.519 after conversion — 100 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.86 → 3.06 (+0.20) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores -0.0 → 1.1631normalised 0.477 → 0.995identity cos to seg 1 0.831 → 0.774 -0.057identity cos neighbours 0.831 → 0.774d_b rescored +0.520 → +0.519d_a rescored +0.520 → +0.519d_a mined 1.163d_b mined 0.518min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_4eZf3luWe6ctotal 19.5schain gain +2.0 dBseam step 0.3 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, good recording, normally alert, slightly relaxed, clear
(thankfulness gratitude · measured, steady, almost no disfluency, formal) biblical foundations, chapter 6, discipleship and leadership, and we are in the supplemental notes.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, slightly guarded; reads as thankfulness gratitude; style: formal, monologue; good recording, no background noise; genuineness 0.5/6; vocal-burst blend 0.2/10; 6.7s, DE.
DE_4eZf3luWe6c_W000000 · in -19.0 dBFS · gain -1.0 dB · emolia-00259
(awe, contentment, affection · slow, fairly steady, little disfluency, storytelling) Baptise them in the name of the Father, the Son and the Holy Spirit. Teaching them to observe all things that I have commanded you. And lo, I am with you even to the end of the age.
full caption & clip details
A middle-aged masculine voice; delivery is normally alert, slow, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, slightly rough, full; clear, little disfluency, wide pitch range, normal breath; affect is mildly positive, slightly dominant, slightly guarded; reads as awe, contentment, affection; style: storytelling, monologue; good recording, quiet background; genuineness 0.2/6; vocal-burst blend 0.0/10; 13.0s, DE.
DE_4eZf3luWe6c_W000004 · in -22.6 dBFS · gain +2.5 dB · emolia-00259
Awe rising ↑identity +0.60 emotion 101 %   sad-Awe-S3-k2 · #13

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Awe.

The raw scorer output across the chain is -0.00, 1.01. In the first clip the scorer found no Awe whatsoever (-0.00); by the last it is at 1.01.

On the corpus-wide percentile scale those become 0.47, 0.99 — a total move of +0.52.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Awe, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Awe sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 12 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.111 before conversion and 0.715 after — it rose by 0.604. Neighbour-to-neighbour the worst pair went 0.111 → 0.715. (The earlier render, with segment 1 left raw, scores 0.423 here.) This chain started below 0.50 — the segments really were different people, the band the conversion helps most: chains starting below 0.50 improve on this measure almost without exception.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Awe moved +0.516 in the original and +0.524 after conversion — 101 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.71 → 2.97 (+0.26) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores -0.0 → 1.0117normalised 0.477 → 0.991identity cos to seg 1 0.111 → 0.715 +0.604identity cos neighbours 0.111 → 0.715d_b rescored +0.516 → +0.524d_a rescored +0.516 → +0.524d_a mined 1.012d_b mined 0.514min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_5CRtHm5ySYgtotal 11.9schain gain +1.7 dBseam step 1.1 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, quiet background, slightly relaxed
(triumph · measured, normally alert, steady, authoritative) Staatliche Hochbaumaßnahmen, und es beginnt der Kollege Norbert Schmitt von der SPD-Fraktion.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, almost no disfluency, fairly narrow pitch, normal breath; affect is neutral, slightly dominant, slightly guarded; reads as triumph; style: authoritative, formal; average recording, quiet background; genuineness 1.9/6; vocal-burst blend 0.3/10; 6.2s, DE.
DE_5CRtHm5ySYg_W000000 · in -20.8 dBFS · gain +0.8 dB · emolia-00125
(awe, contempt, sourness · normal-paced, energised, fairly steady, authoritative) Ich meine, diese Kritik ist schon bemerkenswert, und sie zeigt eine ziemlich erstaunliche Unkenntnis.
full caption & clip details
A middle-aged masculine voice; delivery is energised, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; very clear, little disfluency, moderate pitch range, light breath; affect is neutral, slightly dominant, fairly guarded; reads as awe, contempt, sourness; style: authoritative, formal; good recording, quiet background; genuineness 1.1/6; vocal-burst blend 0.0/10; 6.0s, DE.
DE_5CRtHm5ySYg_W000013 · in -18.0 dBFS · gain -2.0 dB · emolia-00125
Awe rising ↑identity +0.02 emotion 100 %   sad-Awe-S3-k2 · #14

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Awe.

The raw scorer output across the chain is -0.00, 1.49. In the first clip the scorer found no Awe whatsoever (-0.00); by the last it is at 1.49.

On the corpus-wide percentile scale those become 0.47, 1.00 — a total move of +0.52.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Awe, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Awe sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 18 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.809 before conversion and 0.834 after — it rose by 0.025. Neighbour-to-neighbour the worst pair went 0.809 → 0.834. (The earlier render, with segment 1 left raw, scores 0.819 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Awe moved +0.524 in the original and +0.525 after conversion — 100 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.91 → 3.05 (+0.15) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores -0.0 → 1.4863normalised 0.477 → 0.999identity cos to seg 1 0.809 → 0.834 +0.025identity cos neighbours 0.809 → 0.834d_b rescored +0.524 → +0.525d_a rescored +0.524 → +0.525d_a mined 1.486d_b mined 0.522min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_5H96E8-mlsytotal 17.6schain gain +0.8 dBseam step 0.3 dBcrossfades 100 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, no background noise, normal-paced, normally alert, slightly relaxed
(emotional numbness, fear · formal, monologue) riskante Jenseitskontakte medialer Menschen durch Rituale und Intrans.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, fear; style: formal, monologue; good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 4.9s, DE.
DE_5H96E8-mlsy_W000003 · in -12.6 dBFS · gain -7.4 dB · emolia-00229
(awe, infatuation, contemplation · formal, monologue) Deshalb himmlische Wesen zusammen mit dem Gottesgeist die Besitzer und Verwalter von allem sind, was in der himmlischen Schöpfung jemals geschaffen wurde. Was beim Segnen von Geistlichen im Unsichtbaren tatsächlich geschieht.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as awe, infatuation, contemplation; style: formal, monologue; very good recording, no background noise; genuineness 0.1/6; vocal-burst blend 0.0/10; 12.9s, DE.
DE_5H96E8-mlsy_W000006 · in -13.6 dBFS · gain -6.4 dB · emolia-00229
Awe rising ↑identity +0.06 emotion REVERSED   sad-Awe-S3-k2 · #15

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Awe.

The raw scorer output across the chain is -0.00, 1.28. In the first clip the scorer found no Awe whatsoever (-0.00); by the last it is at 1.28.

On the corpus-wide percentile scale those become 0.47, 1.00 — a total move of +0.52.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Awe, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Awe sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 22 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.770 before conversion and 0.834 after — it rose by 0.064. Neighbour-to-neighbour the worst pair went 0.770 → 0.834. (The earlier render, with segment 1 left raw, scores 0.749 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Awe moved +0.522 in the original and +0.000 after conversion — 0 % of the delta retained, so a meaningful part of the trajectory was flattened.

Quality. Mean predicted overall quality across the segments went 2.96 → 3.22 (+0.26) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores -0.0 → 1.2822normalised 0.477 → 0.997identity cos to seg 1 0.770 → 0.834 +0.064identity cos neighbours 0.770 → 0.834d_b rescored +0.522 → +0.000d_a rescored +0.522 → +0.000d_a mined 1.282d_b mined 0.520min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_5NIpRbXP8yktotal 22.0schain gain +1.6 dBseam step 1.0 dBcrossfades 150 ms
Script — 2 chunks, 1 with a non-speech sound
Unchanged across all 2 clips: an adult masculine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, measured, fairly steady, moderate pitch range
(emotional numbness · normally alert, slightly relaxed, some disfluency, monologue) der nun mittlerweile Serie, (ahem) ähm, Drehstrommotor als Generator benutzen. Es hat sich da nämlich was getan. Und zwar ordentlich.
full caption & clip details
An adult masculine voice; delivery is normally alert, measured, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; average clarity, some disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness; style: monologue, didactic; good recording, no background noise; genuineness 1.2/6; vocal-burst blend 0.5/10; 9.0s, DE.
DE_5NIpRbXP8yk_W000000 · in -18.2 dBFS · gain -1.8 dB · emolia-00125
(awe, astonishment surprise, anger · subdued, neutral tension, frequent disfluency, casual) was ich, was ich, wo ich sage, hey, das ist ja echt hammermäßig klasse, ich kann aus einem 160 Watt Drehstrommotor um die 120 Watt rausholen. Das ist, also, poff.
full caption & clip details
A middle-aged masculine voice; delivery is subdued, measured, neutral tension, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; somewhat unclear, frequent disfluency, moderate pitch range, normal breath; affect is neutral, neutral stance, slightly guarded; reads as awe, astonishment surprise, anger; style: casual, monologue; average recording, quiet background; genuineness 4.0/6; vocal-burst blend 0.3/10; 13.2s, DE.
DE_5NIpRbXP8yk_W000047 · in -16.5 dBFS · gain -3.5 dB · emolia-00125
Awe rising ↑identity −0.01 emotion 100 %   sad-Awe-S3-k2 · #16

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Awe.

The raw scorer output across the chain is -0.00, 1.32. In the first clip the scorer found no Awe whatsoever (-0.00); by the last it is at 1.32.

On the corpus-wide percentile scale those become 0.47, 1.00 — a total move of +0.52.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Awe, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Awe sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 16 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.794 before conversion and 0.781 after — it fell by 0.013. Neighbour-to-neighbour the worst pair went 0.794 → 0.781. (The earlier render, with segment 1 left raw, scores 0.725 here.) This chain started 0.70-0.80 — close, but under the identity threshold, a band where the conversion is close to a wash on this measure.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Awe moved +0.523 in the original and +0.524 after conversion — 100 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.66 → 2.91 (+0.26) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores -0.0 → 1.3223normalised 0.477 → 0.998identity cos to seg 1 0.794 → 0.781 -0.013identity cos neighbours 0.794 → 0.781d_b rescored +0.523 → +0.524d_a rescored +0.523 → +0.524d_a mined 1.322d_b mined 0.520min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_5uq5Oz7mtBUtotal 15.2schain gain +1.7 dBseam step 2.6 dBcrossfades 100 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an elderly somewhat feminine voice · slightly warm, dark, smooth, balanced body, very good recording, no background noise, slow, very low-energy
(affection, contentment, contemplation · no disfluency, fairly narrow pitch, whispered, monologue) Es wird dir ganz leicht und du spürst.
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, slow, relaxed, steady; timbre is slightly warm, dark, smooth, balanced body; somewhat unclear, no disfluency, fairly narrow pitch, audible breath; affect is mildly positive, submissive, neutral openness; reads as affection, contentment, contemplation; style: whispered, monologue; very good recording, no background noise; genuineness 1.4/6; vocal-burst blend 6.1/10; 4.1s, DE.
DE_5uq5Oz7mtBU_W000007 · in -16.0 dBFS · gain -4.0 dB · emolia-00062
(contentment, contemplation, awe · little disfluency, narrow pitch range, whispered, monologue) Das dich deine Seele zu einem großen Baum trägt. Es ist der Lebensbaum deiner Vorfahren.
full caption & clip details
An elderly somewhat feminine voice; delivery is very low-energy, slow, relaxed, steady; timbre is slightly warm, dark, smooth, balanced body; somewhat unclear, little disfluency, narrow pitch range, audible breath; affect is mildly positive, submissive, vulnerable; reads as contentment, contemplation, awe; style: whispered, monologue; very good recording, no background noise; genuineness 1.2/6; vocal-burst blend 3.3/10; 11.3s, DE.
DE_5uq5Oz7mtBU_W000008 · in -16.1 dBFS · gain -3.9 dB · emolia-00062
Awe rising ↑identity +0.03 emotion 100 %   sad-Awe-S3-k2 · #17

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Awe.

The raw scorer output across the chain is -0.00, 1.32. In the first clip the scorer found no Awe whatsoever (-0.00); by the last it is at 1.32.

On the corpus-wide percentile scale those become 0.47, 1.00 — a total move of +0.52.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Awe, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Awe sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 29 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.921 before conversion and 0.955 after — it rose by 0.034. Neighbour-to-neighbour the worst pair went 0.921 → 0.955. (The earlier render, with segment 1 left raw, scores 0.848 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Awe moved +0.523 in the original and +0.522 after conversion — 100 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.84 → 3.09 (+0.25) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores -0.0 → 1.3184normalised 0.477 → 0.998identity cos to seg 1 0.921 → 0.955 +0.034identity cos neighbours 0.921 → 0.955d_b rescored +0.523 → +0.522d_a rescored +0.523 → +0.522d_a mined 1.318d_b mined 0.520min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_624drQWANi0total 28.6schain gain +2.8 dBseam step 0.3 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult feminine voice · neutral-toned, neutral-bright, fairly smooth, balanced body, no background noise, normal-paced, normally alert, slightly relaxed
(relief, contentment, distress · almost no disfluency, storytelling, formal) durch das Freiheitszeichen Wassermann lief und die Welt veränderte. Das war eine Zeit des Umsturzes und der Befreiung. Nun steht Pluto wieder an der Schwelle zum Wassermann und auch diesmal wird die Freiheit ein großes Thema sein. Für die Gesellschaft und für den Einzelnen.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, almost no disfluency, moderate pitch range, light breath; affect is mildly positive, slightly dominant, slightly guarded; reads as relief, contentment, distress; style: storytelling, formal; good recording, no background noise; genuineness 0.8/6; vocal-burst blend 0.0/10; 15.3s, DE.
DE_624drQWANi0_W000000 · in -19.7 dBFS · gain -0.3 dB · emolia-00062
(awe, disgust, interest · little disfluency, formal, monologue) Plutus Kräfte der Selbstermächtigung und Erneuerung strömen nun durch euer Wesen. Das ist einerseits aufregend und fantastisch, andererseits ist die Energie aber auch so groß, dass sie verzehrend wirken kann, wenn ihr nicht aufpasst.
full caption & clip details
An adult feminine voice; delivery is normally alert, normal-paced, slightly relaxed, fairly steady; timbre is neutral-toned, neutral-bright, fairly smooth, balanced body; clear, little disfluency, moderate pitch range, light breath; affect is mildly positive, neutral stance, slightly guarded; reads as awe, disgust, interest; style: formal, monologue; very good recording, no background noise; genuineness 0.8/6; vocal-burst blend 0.0/10; 13.5s, DE.
DE_624drQWANi0_W000033 · in -17.2 dBFS · gain -2.8 dB · emolia-00062
Awe rising ↑identity +0.01 emotion 95 %   sad-Awe-S3-k2 · #18

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Awe.

The raw scorer output across the chain is -0.00, 1.00. In the first clip the scorer found no Awe whatsoever (-0.00); by the last it is at 1.00.

On the corpus-wide percentile scale those become 0.47, 0.99 — a total move of +0.52.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Awe, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Awe sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 27 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.908 before conversion and 0.923 after — it rose by 0.015. Neighbour-to-neighbour the worst pair went 0.908 → 0.923. (The earlier render, with segment 1 left raw, scores 0.866 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Awe moved +0.516 in the original and +0.489 after conversion — 95 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.88 → 3.24 (+0.36) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores -0.0 → 1.0029normalised 0.477 → 0.991identity cos to seg 1 0.908 → 0.923 +0.015identity cos neighbours 0.908 → 0.923d_b rescored +0.516 → +0.489d_a rescored +0.516 → +0.489d_a mined 1.003d_b mined 0.514min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_6S01_KQagwMtotal 26.6schain gain +2.0 dBseam step 1.3 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a young adult masculine voice · slightly cool, balanced body, below-average recording, fast, highly aroused, tense, some disfluency, very wide pitch range
(triumph, elation, hope enthusiasm optimism · moderately variable, average clarity, storytelling, dramatic) Und damit hier, Leute, was geht? Und herzlich willkommen zu einer neuen Folge von Newer. Super Mario Bros. Beat. Und wir starten heute rein mit 7-7. Und zwar dem Star Heaven. Und schaut euch dieses Bild an. Ist es nicht wunderschön? Definitiv nein. Und guys, noch eine ganz wichtige Neuigkeit.
full caption & clip details
A young adult masculine voice; delivery is highly aroused, fast, tense, moderately variable; timbre is slightly cool, slightly bright, fairly smooth, balanced body; average clarity, some disfluency, very wide pitch range, normal breath; affect is elated, very dominant, guarded; reads as triumph, elation, hope enthusiasm optimism; style: storytelling, dramatic; below-average recording, noisy background; genuineness 3.4/6; vocal-burst blend 4.4/10; 15.0s, DE.
DE_6S01_KQagwM_W000000 · in -14.1 dBFS · gain -5.9 dB · emolia-00216
(elation, pleasure ecstasy, hope enthusiasm optimism · volatile, very clear, cartoonish, casual) Wie viele Bob-Bombs kommen da? Oh, stellt euch vor, hier wäre einfach schon das Ziel. Das wäre so schön. Eh, das Level war so schwer. Sieht mir verdächtig nach Ziel aus. Let's go.
full caption & clip details
A young adult masculine voice; delivery is highly aroused, fast, tense, volatile; timbre is slightly cool, neutral-bright, very rough, balanced body; very clear, some disfluency, very wide pitch range, normal breath; affect is elated, very dominant, neutral openness; reads as elation, pleasure ecstasy, hope enthusiasm optimism; style: cartoonish, casual; below-average recording, quiet background; genuineness 3.1/6; vocal-burst blend 0.0/10; 11.8s, DE.
DE_6S01_KQagwM_W000024 · in -13.5 dBFS · gain -6.5 dB · emolia-00216
Awe rising ↑identity −0.03 emotion 92 %   sad-Awe-S3-k2 · #19

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Awe.

The raw scorer output across the chain is -0.00, 1.06. In the first clip the scorer found no Awe whatsoever (-0.00); by the last it is at 1.06.

On the corpus-wide percentile scale those become 0.47, 0.99 — a total move of +0.52.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Awe, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Awe sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 12 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 150–150 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.824 before conversion and 0.793 after — it fell by 0.031. Neighbour-to-neighbour the worst pair went 0.824 → 0.793. (The earlier render, with segment 1 left raw, scores 0.673 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Awe moved +0.518 in the original and +0.479 after conversion — 92 % of the delta retained, which is essentially all of it.

Quality. Mean predicted overall quality across the segments went 2.77 → 3.09 (+0.32) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores -0.0 → 1.0635normalised 0.477 → 0.993identity cos to seg 1 0.824 → 0.793 -0.031identity cos neighbours 0.824 → 0.793d_b rescored +0.518 → +0.479d_a rescored +0.518 → +0.479d_a mined 1.063d_b mined 0.516min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_6oWemYVK0_gtotal 11.8schain gain +0.8 dBseam step 0.2 dBcrossfades 150 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: a middle-aged masculine voice · neutral-toned, slightly dark, rough, average recording, quiet background, very low-energy, frequent disfluency, fairly narrow pitch
(contemplation, emotional numbness · slow, slightly relaxed, steady, didactic) Ziel 1 ist Blade in Slovenien. Nummer 2 ist der Seebergsattel.
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, slow, slightly relaxed, steady; timbre is neutral-toned, slightly dark, rough, balanced body; slurred, frequent disfluency, fairly narrow pitch, audible breath; affect is neutral, neutral stance, slightly guarded; reads as contemplation, emotional numbness; style: didactic, whispered; average recording, quiet background; genuineness 1.3/6; vocal-burst blend 0.3/10; 6.8s, DE.
DE_6oWemYVK0_g_W000000 · in -17.4 dBFS · gain -2.6 dB · emolia-00216
(astonishment surprise, awe, anger · measured, relaxed, fairly steady, whispered) Unglaublich. Die Jungs von Österreich wollten sogar unsere Pässe sehen.
full caption & clip details
A middle-aged masculine voice; delivery is very low-energy, measured, relaxed, fairly steady; timbre is neutral-toned, slightly dark, rough, thin; somewhat unclear, frequent disfluency, fairly narrow pitch, audible breath; affect is mildly negative, neutral stance, slightly guarded; reads as astonishment surprise, awe, anger; style: whispered, monologue; average recording, quiet background; genuineness 2.4/6; vocal-burst blend 0.4/10; 5.2s, DE.
DE_6oWemYVK0_g_W000005 · in -20.5 dBFS · gain +0.5 dB · emolia-00216
Awe rising ↑identity −0.00 emotion 3 %   sad-Awe-S3-k2 · #20

This is not a strict-rule trajectory. It comes from rescue rule S3, which exists because the strict rule returns nothing at all for Awe.

The raw scorer output across the chain is -0.00, 1.56. In the first clip the scorer found no Awe whatsoever (-0.00); by the last it is at 1.56.

On the corpus-wide percentile scale those become 0.00, 1.00 — a total move of +1.00.

That is why the strict rule cannot build this chain. Roughly 90 % of the corpus scores exactly zero on Awe, and tied values all collapse onto one point (about 0.45). The clips that genuinely carry Awe sit above 0.85. There is simply nothing in between to step onto, so the only move available is one jump far wider than the 0.25 per-step cap.

What this rule changes: Absent to present. The chain must start at or below 0.05 (the emotion is absent) and end at or above 1.0 (it is clearly present). The most literal reading of 'from not sad to sad'.

What it costs: Says nothing about the SHAPE of the path -- only that it begins absent and ends present. Any intermediate clip is unconstrained.

Same speaker? No similarity score is available here — the emolia clips in this chain are not covered by either speaker-embedding store. The chain therefore rests on the corpus's own speaker/track labelling, which is not the same as a measured check.

Voice consistency: these clips are separate recordings joined together, and no voice-similarity check could be run for this sample, so there is no measurement of how closely the voices match. You may hear the voice shift between segments. Voice conversion has not been applied yet in this build. A planned pass will re-render every segment onto the first segment's voice, which removes this effect entirely.

2 clips · 23 s · emolia

What was done to this chain. All 2 segments were re-synthesised with ChatterboxVC onto segment 1's voice — including segment 1 itself, converted with itself as the target — and then restored with SIDON. Converting the anchor too is what keeps the room and the reverb the same across the whole chain; leaving it raw put a change of acoustic at the first join. The words, timing and delivery still come from each original clip. The joins are 100–100 ms equal-power crossfades. The chain is normalised as one signal, so the loudness differences between segments are the ones the conversion produced, not a per-clip reset.

Did it unify the voice? On the 250-dimensional Speaker-wavLM-id verification embedding, the worst similarity between any segment and segment 1 was 0.945 before conversion and 0.943 after — it fell by 0.002. Neighbour-to-neighbour the worst pair went 0.945 → 0.943. (The earlier render, with segment 1 left raw, scores 0.916 here.) This chain started at or above 0.80 — already effectively one voice. Across the whole build that is the band where conversion tends to cost identity agreement rather than add it, and a hard identity cut would keep this chain without converting it at all. Judge it by ear against the original above.

Did the emotion survive? Re-scored end to end through the same emotion stack and the same corpus-percentile scale the chain was mined on, Awe moved +0.997 in the original and +0.029 after conversion — 3 % of the delta retained, so a meaningful part of the trajectory was flattened.

Quality. Mean predicted overall quality across the segments went 3.02 → 3.15 (+0.13) on the Empathic-Insight head. The corrected render is 48 kHz because SIDON outputs 48 kHz; the original is the 24 kHz source. Some of what you hear as “cleaner” is that bandwidth, not the conversion.

corrected
original
the earlier render — segment 1 left exactly as recorded and only 2…2 converted. Keeps more of the emotional move corpus-wide (96 % vs 90 %) but joins an untouched recording to re-synthesised audio at the first seam, which is what you can hear as the room changing
seg 1 raw
per-clip buttons play:
k 2raw scores -0.0044 → 1.5645normalised 0.002 → 0.999identity cos to seg 1 0.945 → 0.943 -0.002identity cos neighbours 0.945 → 0.943d_b rescored +0.997 → +0.029d_a rescored +0.997 → +0.029d_a mined 1.569d_b mined 0.997min_cos_consec (site) —min_cos_anchor (site) —dataset emolialang despeaker DE_6qdK4WVJHbytotal 23.1schain gain +1.8 dBseam step 0.2 dBcrossfades 100 ms
Script — 2 chunks, 0 with a non-speech sound
Unchanged across all 2 clips: an adult feminine voice · neutral-toned, neutral-bright, fairly smooth, thin, very good recording, no background noise, measured, normally alert
(emotional numbness, sadness, distress · narrow pitch range, formal, monologue) Was geschieht zum Beispiel mit den Körperzellen, wenn ein Mensch aus falsch verstandenem Glauben, seine geheimen Wünsche, wie etwa Zärtlichkeitsgefühle immer wieder verdrängt?
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, no disfluency, narrow pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as emotional numbness, sadness, distress; style: formal, monologue; very good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 12.0s, DE.
DE_6qdK4WVJHby_W000001 · in -13.8 dBFS · gain -6.2 dB · emolia-00099
(awe, emotional numbness, contentment · moderate pitch range, formal, monologue) Als mischgut bezeichnet der Gottesgeist ein geistiges Wissen, das mit weltlichen Speicherungen und den vom Gottesgeist gegebenen Liebebotschaften vermischt ist.
full caption & clip details
An adult feminine voice; delivery is normally alert, measured, slightly relaxed, steady; timbre is neutral-toned, neutral-bright, fairly smooth, thin; clear, no disfluency, moderate pitch range, light breath; affect is neutral, neutral stance, slightly guarded; reads as awe, emotional numbness, contentment; style: formal, monologue; very good recording, no background noise; genuineness 0.0/6; vocal-burst blend 0.0/10; 11.3s, DE.
DE_6qdK4WVJHby_W000018 · in -12.3 dBFS · gain -7.7 dB · emolia-00099