Emotion trajectories, voice-corrected — a parallel build

67 emotion tiers · 1,304 chains · 3,438 segments re-voiced onto their own chain's segment 1 with ChatterboxVC + SIDON. The uncorrected site is untouched and every card here carries both renders.

Read this first. These pages are an experiment with a published result, not a straight upgrade. The emotional trajectories survive the conversion almost intact — that was the risk, and it did not happen. What the conversion does to the voice is not uniform: it is a large improvement on the chains that were genuinely broken and a measurable loss on the chains that were already fine. Both renders are shipped on every card so you can judge by ear rather than take the numbers on trust. The full method and the measurements →
1,304chains converted
3,438segments re-voiced
1.5GPU-hours
0.756 → 0.781median worst-to-anchor identity cosine, processing-matched
96 %median emotion delta retained
The one-paragraph verdict. The emotion survives: re-scored end to end through the same stack and the same corpus-percentile scale the chains were mined on, the trajectory retains a median 90 % of its size in the default render and keeps its direction in 91 % of chains. The voice is a different story, and the average hides it: conversion pulls every chain towards a common level of identity agreement, so what it does depends on where the chain started. For the 25 % whose segments really were different people (cosine below 0.50) it is a large, near-universal win: 0.221 → 0.653. For the 40 % already at or above 0.80 it is roughly neutral at best (-0.015) — and those chains never needed converting. The full breakdown →
Would a hard identity cut be better than converting? For most of the corpus, yes. Of these 1,304 chains, 40 % already clear a 0.80 identity cut, and on those conversion costs 0.117 while helping only 2 % of them. Below 0.70 it is genuinely worth +0.173 — but those are also the chains where you can hear the conversion happen. Section 10 has the thresholds, the corpus-wide counts and the −id/−tbr calibration →
What it costs. Intelligibility is the clearest price: character error rate against the corpus transcript goes 10.0 % → 12.3 % on a stratified sample — the speech survives but is measurably less cleanly articulated.
What each card gives you. The original render and the corrected render side by side, both labelled; a third “uniform” render behind an expander in which segment 1 was converted too; per-chunk playback that follows whichever render you select; the chain's own before/after identity cosine and before/after emotion delta; and everything the original card had — the generated paragraph, the SCRIPT screenplay with its hoisted constants and its (burst) markers, and the full caption behind an expander.
Scope. Emotion tiers only. The VN1 VoiceNet tiers and the vprof_vc voice-profile tiers were deliberately not converted — the profiles are already one cloned voice per chain and have nothing to fix. They are all still on the original site, unchanged.

Manifest tiers — the subsets he will train on

tierrulechainssegments re-voicedcos→seg1 origcorrecteduniformemotion retainedcorrectedoriginal
emotion__B1__T0.20__C0.25__INTERNALB120500.7520.6780.75397 %listen →uncorrected
emotion__B1__T0.25__C0.20__INTERNALB120660.7530.6390.72689 %listen →uncorrected
emotion__B1__T0.25__C0.25__INTERNALB120500.7630.6700.78193 %listen →uncorrected
emotion__B1__T0.40__C0.25__INTERNALB120620.7280.6690.74795 %listen →uncorrected
emotion__B1__T0.50__C0.25__INTERNALB120730.7310.6250.70795 %listen →uncorrected
emotion__B1__T0.60__C0.25__INTERNALB120720.7330.6390.71595 %listen →uncorrected
emotion__B1__T0.70__C0.25__INTERNALB120800.7180.6150.75596 %listen →uncorrected
emotion__B1__T0.80__C0.25__INTERNALB120800.6640.5930.70499 %listen →uncorrected
emotion_twosided__AB2__T0.20__C0.25__INTERNALAB220530.7370.6440.74689 %listen →uncorrected
emotion_twosided__AB2__T0.25__C0.25__INTERNALAB220480.8300.6770.81983 %listen →uncorrected
merged_emo_vn__T0.20__C0.25__INTERNALB1 UNION VN120510.7320.6730.76889 %listen →uncorrected
proxy_spearman__PXR__T0.20__C0.25__INTERNALPXR20440.7740.7010.80295 %listen →uncorrected
proxy_spearman__PXR__T0.25__C0.25__INTERNALPXR20490.7660.6680.78591 %listen →uncorrected
proxy_spearman__PXR__T0.40__C0.25__INTERNALPXR20640.7600.6430.73690 %listen →uncorrected
proxy_spearman__PXR__T0.50__C0.25__INTERNALPXR20740.7760.5780.77099 %listen →uncorrected
proxy_spearman__PXR__T0.60__C0.25__INTERNALPXR20770.6720.5910.68598 %listen →uncorrected
proxy_spearman__PXR__T0.70__C0.25__INTERNALPXR280.4410.5140.66895 %listen →uncorrected
proxy_taillift__PXR__T0.20__C0.25__INTERNALPXR20380.8030.7020.84884 %listen →uncorrected
proxy_taillift__PXR__T0.25__C0.25__INTERNALPXR20540.7480.6890.75089 %listen →uncorrected
proxy_taillift__PXR__T0.40__C0.25__INTERNALPXR20610.7670.6420.75795 %listen →uncorrected
proxy_taillift__PXR__T0.50__C0.25__INTERNALPXR20700.7070.5660.68594 %listen →uncorrected
proxy_taillift__PXR__T0.60__C0.25__INTERNALPXR20760.7090.6010.73598 %listen →uncorrected
proxy_taillift__PXR__T0.70__C0.25__INTERNALPXR280.4410.4840.59597 %listen →uncorrected

Rule × chain length

tierrulechainssegments re-voicedcos→seg1 origcorrecteduniformemotion retainedcorrectedoriginal
k-AB2-k3AB220400.7780.6400.80393 %listen →uncorrected
k-AB2-k4AB220600.7550.7030.78289 %listen →uncorrected
k-AB2-k5AB220800.6750.6320.77198 %listen →uncorrected
k-B1-k2B120200.8410.7230.85476 %listen →uncorrected
k-B1-k3B120400.7950.6960.81194 %listen →uncorrected
k-B1-k4B120600.6610.6340.64992 %listen →uncorrected
k-B1-k5B120800.6640.6140.74595 %listen →uncorrected
k-PXR-k3PXR20400.8350.7190.83592 %listen →uncorrected
k-PXR-k4PXR20600.7830.7060.76099 %listen →uncorrected
k-PXR-k5PXR20800.7300.6940.81398 %listen →uncorrected

One corpus at a time

tierrulechainssegments re-voicedcos→seg1 origcorrecteduniformemotion retainedcorrectedoriginal
c-emolia-AB2AB220530.7870.7010.74989 %listen →uncorrected
c-emolia-B1B120620.6690.6100.73189 %listen →uncorrected
c-emolia-PXRPXR20510.7720.6440.76396 %listen →uncorrected
c-eurospeech-AB2AB220610.4490.6380.83484 %listen →uncorrected
c-eurospeech-B1B120570.6430.7460.87990 %listen →uncorrected
c-eurospeech-PXRPXR20560.3290.6990.86689 %listen →uncorrected
c-mls-AB2AB220570.9260.8060.90496 %listen →uncorrected
c-mls-B1B120440.9300.7860.90599 %listen →uncorrected
c-mls-PXRPXR20570.9170.7740.90898 %listen →uncorrected
c-podcast-AB2AB220500.7010.6580.80594 %listen →uncorrected
c-podcast-B1B120530.5840.5650.73793 %listen →uncorrected
c-podcast-PXRPXR20490.2400.5550.65997 %listen →uncorrected
c-snippets-AB2AB220510.1620.4350.56596 %listen →uncorrected
c-snippets-B1B120610.1790.4400.51793 %listen →uncorrected
c-snippets-PXRPXR20480.1390.4990.61299 %listen →uncorrected
c-evasnippets-AB2AB220640.4820.6630.81297 %listen →uncorrected
c-evasnippets-B1B120590.7690.6590.84699 %listen →uncorrected
c-evasnippets-PXRPXR20590.8630.7650.88899 %listen →uncorrected

Speaker-cleaned set

tierrulechainssegments re-voicedcos→seg1 origcorrecteduniformemotion retainedcorrectedoriginal
sc-AB2-k3AB220400.8440.7330.83394 %listen →uncorrected
sc-AB2-k4AB220600.8350.6890.81994 %listen →uncorrected
sc-AB2-k5AB220800.8250.7290.81797 %listen →uncorrected

Low-resource cells

tierrulechainssegments re-voicedcos→seg1 origcorrecteduniformemotion retainedcorrectedoriginal
rare-AB2-pairsAB220450.6160.5770.70896 %listen →uncorrected
rare-PXR-pairsPXR20550.6200.6590.75599 %listen →uncorrected
rare-B1-pairsB120380.9150.7910.90199 %listen →uncorrected
rare-langmixed20200.9100.7450.860109 %listen →uncorrected

Rescue rules (looser rules, kept separate)

tierrulechainssegments re-voicedcos→seg1 origcorrecteduniformemotion retainedcorrectedoriginal
sad-Sadness-S3-k2S320200.7330.6550.77495 %listen →uncorrected
sad-Sadness-S3-k3S320400.7490.6490.74497 %listen →uncorrected
sad-Sadness-S1-k3S120400.6880.6760.77648 %listen →uncorrected
sad-Sadness-S2-k2S220200.7330.6390.74895 %listen →uncorrected
sad-Sadness-S4-k2S420200.7780.7570.82495 %listen →uncorrected
sad-Awe-S3-k2S320200.7950.7120.79099 %listen →uncorrected
sad-Distress-S3-k2S320200.7180.7020.79593 %listen →uncorrected
sad-Disappointment-S3-k2S320200.7540.7630.81894 %listen →uncorrected
sad-Helplessness-BASE-k3BASE20400.7530.7220.797100 %listen →uncorrected

Provenance

The converted audio is derived from emolia, eurospeech, mls, podcast, snippets and evasnippets. The annotations are CC-BY-4.0. This is a listening demo, not a corpus release, and the converted audio is synthetic speech: it carries the words and delivery of one recording in the voice of another.