Decision continuity under speech compression in voice-pathology detection: A paired audit at matched false-alarm budgets
Background and objective: Telemedicine assesses voice from recordings sent over telephone and video links, where lossy speech codecs commonly sit between the voice and the screening model; clinicians then act on each patient's flag. Codec studies report aggregate accuracy or area under the curve, yet two pipelines with the same false-alarm rate can flag different patients. We ask which patients a codec changes at the operating false-alarm budget, and whether that change exceeds what measuring the same voice again produces. Methods: We audit clean-trained voice-pathology detectors (an acoustic-feature model, three frozen speech-encoder probes and a fine-tuned encoder) on sustained vowels from three public corpora, pairing each speaker's score before and after each of 15 telephony and neural codec settings. At the control-alarm budget of the uncompressed pipeline, each setting's lost detections are compared with a within-vowel yardstick, the loss from re-measuring another stretch of the same vowel without a codec, using a speaker bootstrap and Holm correction across settings. Band-limiting and spectral-shaping shams and level renormalization test alternative explanations, and thresholds recalibrated on separate speakers test whether new thresholds restore the decisions. Results: With the acoustic detector, keeping every previously flagged patient after compression requires flagging more than half of the control speakers for most settings in each of four cohorts. On the largest cohort (1026 speakers, within-corpus training), 12 of 15 settings lose more flagged patients than re-measurement after Holm correction, and the three encoder probes agree (11–13 of 15); for the acoustic detector, band-limiting, spectral shaping and level change do not account for the excess. Recalibrated thresholds keep the false-alarm rate near its target, yet in 10–15 of the 15 settings, depending on the detector, the median loss of previously flagged patients exceeds that of re-measuring the same vowel under the same recalibration. Conclusions: Matching the false-alarm rate after compression does not keep the same patients flagged, so in telemedicine the codec is part of the screening protocol. Before adopting or changing a codec, a remote screening service should report which patients change at its operating budget against a re-measurement yardstick, using only stored paired scores.
Authors
- Von‐Wun Soo (ORCID: https://orcid.org/0000-0002-4810-1244)
- Guan-Yuan Chen (ORCID: https://orcid.org/0000-0003-3298-0624)
Institutions
- Chang Gung University (TW)
- National Tsing Hua University (TW)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-03
- DOI
- https://doi.org/10.5281/zenodo.23117459
- Primary Topic
- Voice and Speech Disorders
- Type
- preprint