Continuous vocal-attribute traversals fall below the path resolution of current speech representations: a pre-registered validity gate across three speaker encoders and three neural codecs
Study 3 of the BedVibe vocal-representation series. Voice conversion, attribute editing and speaker-blend conditioning all move along a straight line between a source and a target point in a learned representation, and treat the intervening coordinates as natural intermediate voices. The path is synthesised, so no ground truth exists anywhere on it except at its two endpoints. This study recorded the path instead. A professionally trained speaker sustains a vowel and slides continuously from clean phonation into rasp, growl or breathiness and back, and from low pitch to high and back — 58 traversals, 2.2–8.5 s each, every intermediate point a real larynx. Two hypotheses and a validity gate with the power to stop the study were pre-registered on 2026-08-14, before any embedding or latent was computed. The gate failed, in both tiers, with zero passes. Across six representations — ECAPA-TDNN, ResNet, WavLM-base-plus-sv, EnCodec, DAC and Mimi — at four window settings and every speaker scope computed, the signal-to-noise ratio of a traversal against the scatter of a steady vowel ranged 0.53–1.85 against a pre-specified requirement of 3.0. A ratio near 1.0 means the displacement produced by a complete vocal transformation is no larger than the displacement observed when the voice is not changing at all. The negative is verified rather than merely observed. Speaker discrimination through the identical code path is positive and rises monotonically with window length in every encoder, so the instrument works. The obvious artifact — phonation attack and decay dominating the endpoint windows — was excluded by re-taking endpoints at 20 % and 80 %, which moved no cell by more than 0.39. The result is a measurement statement about the instruments, and it is quantified: speaker encoders require more audio to produce a stable embedding than a continuous vocal traversal can spare while still resolving its own shape. The two hypotheses are recorded as unreachable with these instruments, not as refuted; no statistic bearing on either was computed. One descriptive observation is reported separately and explicitly outside the hypothesis frame: at 500 ms under ECAPA, a single speaker sliding from clean phonation into rasp traverses a greater distance in the embedding space than the distance separating two different people producing the same vowel in the same condition (cosine 0.28 versus 0.555). No priority or novelty claim is made anywhere in this paper. The audio is not released with this study.
Authors
- Panagiotis Gkilis
Institutions
- Research Studios Austria (AT)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-06
- DOI
- https://doi.org/10.5281/zenodo.22538236
- Primary Topic
- Speech Recognition and Synthesis
- Type
- preprint