Continuous vocal-attribute traversals fall below the path resolution of current speech representations: a pre-registered validity gate across three speaker encoders and three neural codecs

Study 3 of the BedVibe vocal-representation series. Voice conversion, attribute editing and speaker-blend conditioning all move along a straight line between a source and a target point in a learned representation, and treat the intervening coordinates as natural intermediate voices. The path is synthesised, so no ground truth exists anywhere on it except at its two endpoints. This study recorded the path instead. A professionally trained speaker sustains a vowel and slides continuously from clean phonation into rasp, growl or breathiness and back, and from low pitch to high and back — 58 traversals, 2.2–8.5 s each, every intermediate point a real larynx. Two hypotheses and a validity gate with the power to stop the study were pre-registered on 2026-08-14, before any embedding or latent was computed. The gate failed, in both tiers, with zero passes. Across six representations — ECAPA-TDNN, ResNet, WavLM-base-plus-sv, EnCodec, DAC and Mimi — at four window settings and every speaker scope computed, the signal-to-noise ratio of a traversal against the scatter of a steady vowel ranged 0.53–1.85 against a pre-specified requirement of 3.0. A ratio near 1.0 means the displacement produced by a complete vocal transformation is no larger than the displacement observed when the voice is not changing at all. The negative is verified rather than merely observed. Speaker discrimination through the identical code path is positive and rises monotonically with window length in every encoder, so the instrument works. The obvious artifact — phonation attack and decay dominating the endpoint windows — was excluded by re-taking endpoints at 20 % and 80 %, which moved no cell by more than 0.39. The result is a measurement statement about the instruments, and it is quantified: speaker encoders require more audio to produce a stable embedding than a continuous vocal traversal can spare while still resolving its own shape. The two hypotheses are recorded as unreachable with these instruments, not as refuted; no statistic bearing on either was computed. One descriptive observation is reported separately and explicitly outside the hypothesis frame: at 500 ms under ECAPA, a single speaker sliding from clean phonation into rasp traverses a greater distance in the embedding space than the distance separating two different people producing the same vowel in the same condition (cosine 0.28 versus 0.555). No priority or novelty claim is made anywhere in this paper. The audio is not released with this study.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-06
DOI
https://doi.org/10.5281/zenodo.22538236
Primary Topic
Speech Recognition and Synthesis
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Continuous vocal-attribute traversals fall below the path resolution of current speech representations: a pre-registered validity gate across three speaker encoders and three neural codecs

Panagiotis Gkilis
Zenodo (CERN European Organization for Nuclear Research)
Speech Recognition and Synthesis
preprint

Continuous vocal-attribute traversals fall below the path resolution of current speech representations: a pre-registered validity gate across three speaker encoders and three neural codecs

Panagiotis Gkilis
preprint en

Abstract

Study 3 of the BedVibe vocal-representation series. Voice conversion, attribute editing and speaker-blend conditioning all move along a straight line between a source and a target point in a learned representation, and treat the intervening coordinates as natural intermediate voices. The path is synthesised, so no ground truth exists anywhere on it except at its two endpoints. This study recorded the path instead. A professionally trained speaker sustains a vowel and slides continuously from clean phonation into rasp, growl or breathiness and back, and from low pitch to high and back — 58 traversals, 2.2–8.5 s each, every intermediate point a real larynx. Two hypotheses and a validity gate with the power to stop the study were pre-registered on 2026-08-14, before any embedding or latent was computed. The gate failed, in both tiers, with zero passes. Across six representations — ECAPA-TDNN, ResNet, WavLM-base-plus-sv, EnCodec, DAC and Mimi — at four window settings and every speaker scope computed, the signal-to-noise ratio of a traversal against the scatter of a steady vowel ranged 0.53–1.85 against a pre-specified requirement of 3.0. A ratio near 1.0 means the displacement produced by a complete vocal transformation is no larger than the displacement observed when the voice is not changing at all. The negative is verified rather than merely observed. Speaker discrimination through the identical code path is positive and rises monotonically with window length in every encoder, so the instrument works. The obvious artifact — phonation attack and decay dominating the endpoint windows — was excluded by re-taking endpoints at 20 % and 80 %, which moved no cell by more than 0.39. The result is a measurement statement about the instruments, and it is quantified: speaker encoders require more audio to produce a stable embedding than a continuous vocal traversal can spare while still resolving its own shape. The two hypotheses are recorded as unreachable with these instruments, not as refuted; no statistic bearing on either was computed. One descriptive observation is reported separately and explicitly outside the hypothesis frame: at 500 ms under ECAPA, a single speaker sliding from clean phonation into rasp traverses a greater distance in the embedding space than the distance separating two different people producing the same vowel in the same condition (cosine 0.28 versus 0.555). No priority or novelty claim is made anywhere in this paper. The audio is not released with this study.

Zenodo (CERN European Organization for Nuclear Research)
Research Studios Austria (AT)
Peace, Justice and strong institutions, Reduced inequalities
Speech Recognition and Synthesis
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.