Decision continuity under speech compression in voice-pathology detection: A paired audit at matched false-alarm budgets

Background and objective: Telemedicine assesses voice from recordings sent over telephone and video links, where lossy speech codecs commonly sit between the voice and the screening model, and clinicians act on each patient’s flag. Codec studies report aggregate accuracy, yet two pipelines with the same false-alarm rate can flag different patients. We ask which patients a codec changes at the operating false-alarm budget, and whether this exceeds re-measuring the same voice. Methods: We audit clean-trained voice-pathology detectors (an acoustic-feature model, three frozen speech-encoder probes, a fine-tuned encoder) on sustained vowels from three public corpora, pairing each speaker’s score before and after each of 15 telephony and neural codec settings. At the uncompressed pipeline’s control-alarm budget, each setting’s lost detections are compared with a within-vowel yardstick, the loss from re-measuring another stretch of the same vowel without a codec, using a speaker bootstrap with Holm correction across settings. Band-limiting and spectral-shaping shams and level renormalization test alternative explanations, and recalibration on separate speakers tests a remedy. Results: On the largest cohort (1026 speakers, within-corpus training), compression changes which patients are flagged even where sensitivity rises: with AMR-WB, the acoustic detector’s sensitivity rises from 35.6% to 40.5%, yet 29 previously flagged patients are lost (on vowel halves, 49 against 28 for re-measurement) and 52 others newly flagged. Of the 15 settings studied, 12 lose more flagged patients than re-measurement (9–12 over six cross-validation fold assignments), and three encoder probes agree (11–13 of 15); for the acoustic detector, band-limiting, spectral shaping and level do not explain the excess. For cross-corpus-trained detectors, recalibration keeps the median false-alarm rate over splits near its target, yet in 10–15 of 15 settings, depending on the detector, its median loss of previously flagged patients exceeds that of equally recalibrated re-measurement. Conclusions: On the largest cohort, most settings studied lose more flagged patients than re-measurement at a matched false-alarm rate, so in telemedicine the codec is part of the screening protocol. Before adopting or changing a codec, a screening service should report which patients change at its operating budget against a re-measurement yardstick; stored paired scores suffice.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-05
DOI
https://doi.org/10.5281/zenodo.23153791
Primary Topic
Voice and Speech Disorders
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Decision continuity under speech compression in voice-pathology detection: A paired audit at matched false-alarm budgets

Von‐Wun Soo, Guan-Yuan Chen
Zenodo (CERN European Organization for Nuclear Research)
Voice and Speech Disorders
preprint

Decision continuity under speech compression in voice-pathology detection: A paired audit at matched false-alarm budgets

Von‐Wun Soo, Guan-Yuan Chen
preprint en

Abstract

Background and objective: Telemedicine assesses voice from recordings sent over telephone and video links, where lossy speech codecs commonly sit between the voice and the screening model, and clinicians act on each patient’s flag. Codec studies report aggregate accuracy, yet two pipelines with the same false-alarm rate can flag different patients. We ask which patients a codec changes at the operating false-alarm budget, and whether this exceeds re-measuring the same voice. Methods: We audit clean-trained voice-pathology detectors (an acoustic-feature model, three frozen speech-encoder probes, a fine-tuned encoder) on sustained vowels from three public corpora, pairing each speaker’s score before and after each of 15 telephony and neural codec settings. At the uncompressed pipeline’s control-alarm budget, each setting’s lost detections are compared with a within-vowel yardstick, the loss from re-measuring another stretch of the same vowel without a codec, using a speaker bootstrap with Holm correction across settings. Band-limiting and spectral-shaping shams and level renormalization test alternative explanations, and recalibration on separate speakers tests a remedy. Results: On the largest cohort (1026 speakers, within-corpus training), compression changes which patients are flagged even where sensitivity rises: with AMR-WB, the acoustic detector’s sensitivity rises from 35.6% to 40.5%, yet 29 previously flagged patients are lost (on vowel halves, 49 against 28 for re-measurement) and 52 others newly flagged. Of the 15 settings studied, 12 lose more flagged patients than re-measurement (9–12 over six cross-validation fold assignments), and three encoder probes agree (11–13 of 15); for the acoustic detector, band-limiting, spectral shaping and level do not explain the excess. For cross-corpus-trained detectors, recalibration keeps the median false-alarm rate over splits near its target, yet in 10–15 of 15 settings, depending on the detector, its median loss of previously flagged patients exceeds that of equally recalibrated re-measurement. Conclusions: On the largest cohort, most settings studied lose more flagged patients than re-measurement at a matched false-alarm rate, so in telemedicine the codec is part of the screening protocol. Before adopting or changing a codec, a screening service should report which patients change at its operating budget against a re-measurement yardstick; stored paired scores suffice.

Zenodo (CERN European Organization for Nuclear Research)
Chang Gung University (TW), National Tsing Hua University (TW)
Voice and Speech Disorders
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.