Class-imbalance-sensitive evaluation of automated decisions in mixed heart-lung auscultation
OBJECTIVE: Automated auscultation models are commonly evaluated using aggregate metrics, although mixed heart-lung recordings may be severely imbalanced and contain related acquisition conditions. This study examined whether prevalence, grouping, imbalance correction and decision-threshold selection materially affect the interpretation of cardiac and pulmonary abnormality decisions. Approach: A deterministic log-mel pipeline and nested grouped evaluation based on metadata-defined heart-lung sound-state combinations were applied to 145 mixed HLS-CMDS manikin recordings forming 60 operational groups. Independent and shared convolutional neural networks (CNNs) and a logistic-summary comparator were assessed using five repeated six-fold outer evaluations, with thresholds selected exclusively from inner grouped out-of-fold predictions. All-positive baselines, imbalance-training ablations, preprocessing sensitivity and support requirements for context analysis were evaluated. Results: The deterministic all-positive baselines achieved F1-scores of 0.953 for cardiac detection and 0.893 for pulmonary detection while failing to identify any normal recordings, demonstrating that high F1-score could be prevalence-driven. Across the five repeat-pooled outer evaluations, the logistic-summary comparator achieved the highest pulmonary F1-score, balanced accuracy and Matthews correlation coefficient (MCC), with mean ± SD values of 0.904 ± 0.015, 0.703 ± 0.041 and 0.451 ± 0.077, respectively. The independent CNN achieved the highest pulmonary specificity of 0.643 ± 0.143. Cardiac normal-class recovery remained unstable and did not support robust discrimination. For the independent pulmonary CNN, combined class weighting and oversampling did not improve balanced accuracy or MCC relative to class weighting alone. Support-aware analysis identified insufficient normal-context support for a reliable between-context comparison. Significance: Grouped partitioning, trivial baselines, nested threshold selection and normal-class-sensitive metrics materially changed performance interpretation. The results support methodological safeguards for automated auscultation evaluation but do not establish clinical diagnostic validity. .
Authors
- Lulu Wang (ORCID: https://orcid.org/0000-0001-7466-9522)
Institutions
- Reykjavík University (IS)
Publication Details
- Journal
- Physiological Measurement
- Published
- 2026-09-18
- DOI
- https://doi.org/10.1088/1361-6579/aea9eb
- Primary Topic
- Phonocardiography and Auscultation Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00