Decision continuity under speech compression in voice-pathology detection: A paired audit at matched false-alarm budgets
Background and objective: Telemedicine assesses voice from recordings sent over telephone and video links, where lossy speech codecs commonly sit between the voice and the screening model, and clinicians act on each patient’s flag. Codec studies report aggregate accuracy, yet two pipelines with the same false-alarm rate can flag different patients. We ask which patients a codec changes at the operating false-alarm budget, and whether this exceeds re-measuring the same voice. Methods: We audit clean-trained voice-pathology detectors (an acoustic-feature model, three frozen speech-encoder probes, a fine-tuned encoder) on sustained vowels from three public corpora, pairing each speaker’s score before and after each of 15 telephony and neural codec settings. At the uncompressed pipeline’s control-alarm budget, each setting’s lost detections are compared with a within-vowel yardstick, the loss from re-measuring another stretch of the same vowel without a codec, using a speaker bootstrap with Holm correction across settings. Band-limiting and spectral-shaping shams and level renormalization test alternative explanations, and recalibration on separate speakers tests a remedy. Results: On the largest cohort (1026 speakers, within-corpus training), compression changes which patients are flagged even where sensitivity rises: with AMR-WB, the acoustic detector’s sensitivity rises from 35.6% to 40.5%, yet 29 previously flagged patients are lost (on vowel halves, 49 against 28 for re-measurement) and 52 others newly flagged. Of the 15 settings studied, 12 lose more flagged patients than re-measurement (9–12 over six cross-validation fold assignments), and three encoder probes agree (11–13 of 15); for the acoustic detector, band-limiting, spectral shaping and level do not explain the excess. For cross-corpus-trained detectors, recalibration keeps the median false-alarm rate over splits near its target, yet in 10–15 of 15 settings, depending on the detector, its median loss of previously flagged patients exceeds that of equally recalibrated re-measurement. Conclusions: On the largest cohort, most settings studied lose more flagged patients than re-measurement at a matched false-alarm rate, so in telemedicine the codec is part of the screening protocol. Before adopting or changing a codec, a screening service should report which patients change at its operating budget against a re-measurement yardstick; stored paired scores suffice.
Authors
- Von‐Wun Soo (ORCID: https://orcid.org/0000-0002-4810-1244)
- Guan-Yuan Chen (ORCID: https://orcid.org/0000-0003-3298-0624)
Institutions
- Chang Gung University (TW)
- National Tsing Hua University (TW)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-05
- DOI
- https://doi.org/10.5281/zenodo.23153791
- Primary Topic
- Voice and Speech Disorders
- Type
- preprint