Decision continuity under speech compression in voice-pathology detection: A paired audit at matched false-alarm budgets

Background and objective: Telemedicine assesses voice from recordings sent over telephone and video links, where lossy speech codecs commonly sit between the voice and the screening model; clinicians then act on each patient's flag. Codec studies report aggregate accuracy or area under the curve, yet two pipelines with the same false-alarm rate can flag different patients. We ask which patients a codec changes at the operating false-alarm budget, and whether that change exceeds what measuring the same voice again produces. Methods: We audit clean-trained voice-pathology detectors (an acoustic-feature model, three frozen speech-encoder probes and a fine-tuned encoder) on sustained vowels from three public corpora, pairing each speaker's score before and after each of 15 telephony and neural codec settings. At the control-alarm budget of the uncompressed pipeline, each setting's lost detections are compared with a within-vowel yardstick, the loss from re-measuring another stretch of the same vowel without a codec, using a speaker bootstrap and Holm correction across settings. Band-limiting and spectral-shaping shams and level renormalization test alternative explanations, and thresholds recalibrated on separate speakers test whether new thresholds restore the decisions. Results: With the acoustic detector, keeping every previously flagged patient after compression requires flagging more than half of the control speakers for most settings in each of four cohorts. On the largest cohort (1026 speakers, within-corpus training), 12 of 15 settings lose more flagged patients than re-measurement after Holm correction, and the three encoder probes agree (11–13 of 15); for the acoustic detector, band-limiting, spectral shaping and level change do not account for the excess. Recalibrated thresholds keep the false-alarm rate near its target, yet in 10–15 of the 15 settings, depending on the detector, the median loss of previously flagged patients exceeds that of re-measuring the same vowel under the same recalibration. Conclusions: Matching the false-alarm rate after compression does not keep the same patients flagged, so in telemedicine the codec is part of the screening protocol. Before adopting or changing a codec, a remote screening service should report which patients change at its operating budget against a re-measurement yardstick, using only stored paired scores.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-03
DOI
https://doi.org/10.5281/zenodo.23117459
Primary Topic
Voice and Speech Disorders
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Decision continuity under speech compression in voice-pathology detection: A paired audit at matched false-alarm budgets

Von‐Wun Soo, Guan-Yuan Chen
Zenodo (CERN European Organization for Nuclear Research)
Voice and Speech Disorders
preprint

Decision continuity under speech compression in voice-pathology detection: A paired audit at matched false-alarm budgets

Von‐Wun Soo, Guan-Yuan Chen
preprint en

Abstract

Background and objective: Telemedicine assesses voice from recordings sent over telephone and video links, where lossy speech codecs commonly sit between the voice and the screening model; clinicians then act on each patient's flag. Codec studies report aggregate accuracy or area under the curve, yet two pipelines with the same false-alarm rate can flag different patients. We ask which patients a codec changes at the operating false-alarm budget, and whether that change exceeds what measuring the same voice again produces. Methods: We audit clean-trained voice-pathology detectors (an acoustic-feature model, three frozen speech-encoder probes and a fine-tuned encoder) on sustained vowels from three public corpora, pairing each speaker's score before and after each of 15 telephony and neural codec settings. At the control-alarm budget of the uncompressed pipeline, each setting's lost detections are compared with a within-vowel yardstick, the loss from re-measuring another stretch of the same vowel without a codec, using a speaker bootstrap and Holm correction across settings. Band-limiting and spectral-shaping shams and level renormalization test alternative explanations, and thresholds recalibrated on separate speakers test whether new thresholds restore the decisions. Results: With the acoustic detector, keeping every previously flagged patient after compression requires flagging more than half of the control speakers for most settings in each of four cohorts. On the largest cohort (1026 speakers, within-corpus training), 12 of 15 settings lose more flagged patients than re-measurement after Holm correction, and the three encoder probes agree (11–13 of 15); for the acoustic detector, band-limiting, spectral shaping and level change do not account for the excess. Recalibrated thresholds keep the false-alarm rate near its target, yet in 10–15 of the 15 settings, depending on the detector, the median loss of previously flagged patients exceeds that of re-measuring the same vowel under the same recalibration. Conclusions: Matching the false-alarm rate after compression does not keep the same patients flagged, so in telemedicine the codec is part of the screening protocol. Before adopting or changing a codec, a remote screening service should report which patients change at its operating budget against a re-measurement yardstick, using only stored paired scores.

Zenodo (CERN European Organization for Nuclear Research)
Chang Gung University (TW), National Tsing Hua University (TW)
Voice and Speech Disorders
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.