Decision continuity under speech compression in voice-pathology detection: A paired audit at matched false-alarm budgets
Background and objective: Remote voice assessment increasingly receives speech that has passed through a codec, whereas voice-pathology detectors are usually developed on uncompressed recordings. Accuracy after compression does not show whether previously flagged patients are still flagged: exchanged detections can leave sensitivity unchanged, and keeping the original (inherited) threshold can preserve detections only by raising false alarms. We aimed to measure this decision continuity at a fixed false-alarm cost and to compare it with within-vowel variation. Methods: We audit fixed scoring pipelines with paired case identifiers: each transformed score set is cut at no more than the reference number of control alarms, and order statistics count retained and lost reference detections and newly detected positive cases. We apply the audit to 15 codec settings on sustained vowels (PVQD, SVD, VOICED) with acoustic and WavLM detectors, compare codec losses with a no-codec within-vowel reference (the other half of the same phonation; a pre-specified contrast and a post hoc cross-stretch contrast with matched stretch structure), and transfer calibrated thresholds to held-out SVD speakers with a PVQD-trained model. Results: On matched PVQD, 5 codec settings lose no detection at the inherited threshold but raise false-positive rates to 67.9–100%. Against the within-vowel reference, FocalCodec (three frame rates), X-Codec 2 and EnCodec show a mean excess of 8–15.5 lost reference detections across swapped-half comparisons on matched PVQD (not additional patients), with ranges above zero under both contrasts and in both PVQD cohorts; Opus exceeds it under the cross-stretch contrast. For the other nine codecs, including Mimi and X-Codec, excess was not resolved on matched PVQD: their descriptive ranges include zero under both contrasts (three exceed on full PVQD under the cross-stretch contrast only). With transferred thresholds, median loss is 8.3–86.0% at near-target median test false-positive rates. Conclusions: Paired, budget-matched reporting shows which detections change after a channel change and at what false-alarm cost, which aggregate accuracy and inherited-threshold retention do not. Read beside a no-codec, other-half reference, it separates settings with excess turnover from those for which descriptive ranges do not resolve excess. It needs only stored paired scores and sorting.
Authors
- Von‐Wun Soo (ORCID: https://orcid.org/0000-0002-4810-1244)
- Guan-Yuan Chen (ORCID: https://orcid.org/0000-0003-3298-0624)
Institutions
- Chang Gung University (TW)
- National Tsing Hua University (TW)
- North Carolina Exploring Cultural Heritage Online (US)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-24
- DOI
- https://doi.org/10.5281/zenodo.22675321
- Primary Topic
- Voice and Speech Disorders
- Type
- preprint