67 emotion tiers · 1,304 chains · 3,438 segments re-voiced onto their own chain's segment 1 with ChatterboxVC + SIDON. The uncorrected site is untouched and every card here carries both renders.
VN1 VoiceNet tiers and the vprof_vc voice-profile tiers were
deliberately not converted — the profiles are already one cloned voice per chain
and have nothing to fix. They are all still on
the original site, unchanged.| tier | rule | chains | segments re-voiced | cos→seg1 orig | corrected | uniform | emotion retained | corrected | original |
|---|---|---|---|---|---|---|---|---|---|
emotion__B1__T0.20__C0.25__INTERNAL | B1 | 20 | 50 | 0.752 | 0.678 | 0.753 | 97 % | listen → | uncorrected |
emotion__B1__T0.25__C0.20__INTERNAL | B1 | 20 | 66 | 0.753 | 0.639 | 0.726 | 89 % | listen → | uncorrected |
emotion__B1__T0.25__C0.25__INTERNAL | B1 | 20 | 50 | 0.763 | 0.670 | 0.781 | 93 % | listen → | uncorrected |
emotion__B1__T0.40__C0.25__INTERNAL | B1 | 20 | 62 | 0.728 | 0.669 | 0.747 | 95 % | listen → | uncorrected |
emotion__B1__T0.50__C0.25__INTERNAL | B1 | 20 | 73 | 0.731 | 0.625 | 0.707 | 95 % | listen → | uncorrected |
emotion__B1__T0.60__C0.25__INTERNAL | B1 | 20 | 72 | 0.733 | 0.639 | 0.715 | 95 % | listen → | uncorrected |
emotion__B1__T0.70__C0.25__INTERNAL | B1 | 20 | 80 | 0.718 | 0.615 | 0.755 | 96 % | listen → | uncorrected |
emotion__B1__T0.80__C0.25__INTERNAL | B1 | 20 | 80 | 0.664 | 0.593 | 0.704 | 99 % | listen → | uncorrected |
emotion_twosided__AB2__T0.20__C0.25__INTERNAL | AB2 | 20 | 53 | 0.737 | 0.644 | 0.746 | 89 % | listen → | uncorrected |
emotion_twosided__AB2__T0.25__C0.25__INTERNAL | AB2 | 20 | 48 | 0.830 | 0.677 | 0.819 | 83 % | listen → | uncorrected |
merged_emo_vn__T0.20__C0.25__INTERNAL | B1 UNION VN1 | 20 | 51 | 0.732 | 0.673 | 0.768 | 89 % | listen → | uncorrected |
proxy_spearman__PXR__T0.20__C0.25__INTERNAL | PXR | 20 | 44 | 0.774 | 0.701 | 0.802 | 95 % | listen → | uncorrected |
proxy_spearman__PXR__T0.25__C0.25__INTERNAL | PXR | 20 | 49 | 0.766 | 0.668 | 0.785 | 91 % | listen → | uncorrected |
proxy_spearman__PXR__T0.40__C0.25__INTERNAL | PXR | 20 | 64 | 0.760 | 0.643 | 0.736 | 90 % | listen → | uncorrected |
proxy_spearman__PXR__T0.50__C0.25__INTERNAL | PXR | 20 | 74 | 0.776 | 0.578 | 0.770 | 99 % | listen → | uncorrected |
proxy_spearman__PXR__T0.60__C0.25__INTERNAL | PXR | 20 | 77 | 0.672 | 0.591 | 0.685 | 98 % | listen → | uncorrected |
proxy_spearman__PXR__T0.70__C0.25__INTERNAL | PXR | 2 | 8 | 0.441 | 0.514 | 0.668 | 95 % | listen → | uncorrected |
proxy_taillift__PXR__T0.20__C0.25__INTERNAL | PXR | 20 | 38 | 0.803 | 0.702 | 0.848 | 84 % | listen → | uncorrected |
proxy_taillift__PXR__T0.25__C0.25__INTERNAL | PXR | 20 | 54 | 0.748 | 0.689 | 0.750 | 89 % | listen → | uncorrected |
proxy_taillift__PXR__T0.40__C0.25__INTERNAL | PXR | 20 | 61 | 0.767 | 0.642 | 0.757 | 95 % | listen → | uncorrected |
proxy_taillift__PXR__T0.50__C0.25__INTERNAL | PXR | 20 | 70 | 0.707 | 0.566 | 0.685 | 94 % | listen → | uncorrected |
proxy_taillift__PXR__T0.60__C0.25__INTERNAL | PXR | 20 | 76 | 0.709 | 0.601 | 0.735 | 98 % | listen → | uncorrected |
proxy_taillift__PXR__T0.70__C0.25__INTERNAL | PXR | 2 | 8 | 0.441 | 0.484 | 0.595 | 97 % | listen → | uncorrected |
| tier | rule | chains | segments re-voiced | cos→seg1 orig | corrected | uniform | emotion retained | corrected | original |
|---|---|---|---|---|---|---|---|---|---|
k-AB2-k3 | AB2 | 20 | 40 | 0.778 | 0.640 | 0.803 | 93 % | listen → | uncorrected |
k-AB2-k4 | AB2 | 20 | 60 | 0.755 | 0.703 | 0.782 | 89 % | listen → | uncorrected |
k-AB2-k5 | AB2 | 20 | 80 | 0.675 | 0.632 | 0.771 | 98 % | listen → | uncorrected |
k-B1-k2 | B1 | 20 | 20 | 0.841 | 0.723 | 0.854 | 76 % | listen → | uncorrected |
k-B1-k3 | B1 | 20 | 40 | 0.795 | 0.696 | 0.811 | 94 % | listen → | uncorrected |
k-B1-k4 | B1 | 20 | 60 | 0.661 | 0.634 | 0.649 | 92 % | listen → | uncorrected |
k-B1-k5 | B1 | 20 | 80 | 0.664 | 0.614 | 0.745 | 95 % | listen → | uncorrected |
k-PXR-k3 | PXR | 20 | 40 | 0.835 | 0.719 | 0.835 | 92 % | listen → | uncorrected |
k-PXR-k4 | PXR | 20 | 60 | 0.783 | 0.706 | 0.760 | 99 % | listen → | uncorrected |
k-PXR-k5 | PXR | 20 | 80 | 0.730 | 0.694 | 0.813 | 98 % | listen → | uncorrected |
| tier | rule | chains | segments re-voiced | cos→seg1 orig | corrected | uniform | emotion retained | corrected | original |
|---|---|---|---|---|---|---|---|---|---|
c-emolia-AB2 | AB2 | 20 | 53 | 0.787 | 0.701 | 0.749 | 89 % | listen → | uncorrected |
c-emolia-B1 | B1 | 20 | 62 | 0.669 | 0.610 | 0.731 | 89 % | listen → | uncorrected |
c-emolia-PXR | PXR | 20 | 51 | 0.772 | 0.644 | 0.763 | 96 % | listen → | uncorrected |
c-eurospeech-AB2 | AB2 | 20 | 61 | 0.449 | 0.638 | 0.834 | 84 % | listen → | uncorrected |
c-eurospeech-B1 | B1 | 20 | 57 | 0.643 | 0.746 | 0.879 | 90 % | listen → | uncorrected |
c-eurospeech-PXR | PXR | 20 | 56 | 0.329 | 0.699 | 0.866 | 89 % | listen → | uncorrected |
c-mls-AB2 | AB2 | 20 | 57 | 0.926 | 0.806 | 0.904 | 96 % | listen → | uncorrected |
c-mls-B1 | B1 | 20 | 44 | 0.930 | 0.786 | 0.905 | 99 % | listen → | uncorrected |
c-mls-PXR | PXR | 20 | 57 | 0.917 | 0.774 | 0.908 | 98 % | listen → | uncorrected |
c-podcast-AB2 | AB2 | 20 | 50 | 0.701 | 0.658 | 0.805 | 94 % | listen → | uncorrected |
c-podcast-B1 | B1 | 20 | 53 | 0.584 | 0.565 | 0.737 | 93 % | listen → | uncorrected |
c-podcast-PXR | PXR | 20 | 49 | 0.240 | 0.555 | 0.659 | 97 % | listen → | uncorrected |
c-snippets-AB2 | AB2 | 20 | 51 | 0.162 | 0.435 | 0.565 | 96 % | listen → | uncorrected |
c-snippets-B1 | B1 | 20 | 61 | 0.179 | 0.440 | 0.517 | 93 % | listen → | uncorrected |
c-snippets-PXR | PXR | 20 | 48 | 0.139 | 0.499 | 0.612 | 99 % | listen → | uncorrected |
c-evasnippets-AB2 | AB2 | 20 | 64 | 0.482 | 0.663 | 0.812 | 97 % | listen → | uncorrected |
c-evasnippets-B1 | B1 | 20 | 59 | 0.769 | 0.659 | 0.846 | 99 % | listen → | uncorrected |
c-evasnippets-PXR | PXR | 20 | 59 | 0.863 | 0.765 | 0.888 | 99 % | listen → | uncorrected |
| tier | rule | chains | segments re-voiced | cos→seg1 orig | corrected | uniform | emotion retained | corrected | original |
|---|---|---|---|---|---|---|---|---|---|
sc-AB2-k3 | AB2 | 20 | 40 | 0.844 | 0.733 | 0.833 | 94 % | listen → | uncorrected |
sc-AB2-k4 | AB2 | 20 | 60 | 0.835 | 0.689 | 0.819 | 94 % | listen → | uncorrected |
sc-AB2-k5 | AB2 | 20 | 80 | 0.825 | 0.729 | 0.817 | 97 % | listen → | uncorrected |
| tier | rule | chains | segments re-voiced | cos→seg1 orig | corrected | uniform | emotion retained | corrected | original |
|---|---|---|---|---|---|---|---|---|---|
rare-AB2-pairs | AB2 | 20 | 45 | 0.616 | 0.577 | 0.708 | 96 % | listen → | uncorrected |
rare-PXR-pairs | PXR | 20 | 55 | 0.620 | 0.659 | 0.755 | 99 % | listen → | uncorrected |
rare-B1-pairs | B1 | 20 | 38 | 0.915 | 0.791 | 0.901 | 99 % | listen → | uncorrected |
rare-lang | mixed | 20 | 20 | 0.910 | 0.745 | 0.860 | 109 % | listen → | uncorrected |
| tier | rule | chains | segments re-voiced | cos→seg1 orig | corrected | uniform | emotion retained | corrected | original |
|---|---|---|---|---|---|---|---|---|---|
sad-Sadness-S3-k2 | S3 | 20 | 20 | 0.733 | 0.655 | 0.774 | 95 % | listen → | uncorrected |
sad-Sadness-S3-k3 | S3 | 20 | 40 | 0.749 | 0.649 | 0.744 | 97 % | listen → | uncorrected |
sad-Sadness-S1-k3 | S1 | 20 | 40 | 0.688 | 0.676 | 0.776 | 48 % | listen → | uncorrected |
sad-Sadness-S2-k2 | S2 | 20 | 20 | 0.733 | 0.639 | 0.748 | 95 % | listen → | uncorrected |
sad-Sadness-S4-k2 | S4 | 20 | 20 | 0.778 | 0.757 | 0.824 | 95 % | listen → | uncorrected |
sad-Awe-S3-k2 | S3 | 20 | 20 | 0.795 | 0.712 | 0.790 | 99 % | listen → | uncorrected |
sad-Distress-S3-k2 | S3 | 20 | 20 | 0.718 | 0.702 | 0.795 | 93 % | listen → | uncorrected |
sad-Disappointment-S3-k2 | S3 | 20 | 20 | 0.754 | 0.763 | 0.818 | 94 % | listen → | uncorrected |
sad-Helplessness-BASE-k3 | BASE | 20 | 40 | 0.753 | 0.722 | 0.797 | 100 % | listen → | uncorrected |
emolia, eurospeech, mls, podcast,
snippets and evasnippets. The annotations are CC-BY-4.0. This
is a listening demo, not a corpus release, and the converted audio is synthetic
speech: it carries the words and delivery of one recording in the voice of
another.