Class-imbalance-sensitive evaluation of automated decisions in mixed heart-lung auscultation

OBJECTIVE: Automated auscultation models are commonly evaluated using aggregate metrics, although mixed heart-lung recordings may be severely imbalanced and contain related acquisition conditions. This study examined whether prevalence, grouping, imbalance correction and decision-threshold selection materially affect the interpretation of cardiac and pulmonary abnormality decisions. Approach: A deterministic log-mel pipeline and nested grouped evaluation based on metadata-defined heart-lung sound-state combinations were applied to 145 mixed HLS-CMDS manikin recordings forming 60 operational groups. Independent and shared convolutional neural networks (CNNs) and a logistic-summary comparator were assessed using five repeated six-fold outer evaluations, with thresholds selected exclusively from inner grouped out-of-fold predictions. All-positive baselines, imbalance-training ablations, preprocessing sensitivity and support requirements for context analysis were evaluated. Results: The deterministic all-positive baselines achieved F1-scores of 0.953 for cardiac detection and 0.893 for pulmonary detection while failing to identify any normal recordings, demonstrating that high F1-score could be prevalence-driven. Across the five repeat-pooled outer evaluations, the logistic-summary comparator achieved the highest pulmonary F1-score, balanced accuracy and Matthews correlation coefficient (MCC), with mean ± SD values of 0.904 ± 0.015, 0.703 ± 0.041 and 0.451 ± 0.077, respectively. The independent CNN achieved the highest pulmonary specificity of 0.643 ± 0.143. Cardiac normal-class recovery remained unstable and did not support robust discrimination. For the independent pulmonary CNN, combined class weighting and oversampling did not improve balanced accuracy or MCC relative to class weighting alone. Support-aware analysis identified insufficient normal-context support for a reliable between-context comparison. Significance: Grouped partitioning, trivial baselines, nested threshold selection and normal-class-sensitive metrics materially changed performance interpretation. The results support methodological safeguards for automated auscultation evaluation but do not establish clinical diagnostic validity. .

Authors

Institutions

Publication Details

Journal
Physiological Measurement
Published
2026-09-18
DOI
https://doi.org/10.1088/1361-6579/aea9eb
Primary Topic
Phonocardiography and Auscultation Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Class-imbalance-sensitive evaluation of automated decisions in mixed heart-lung auscultation

Lulu Wang
Physiological Measurement
Phonocardiography and Auscultation Techniques
article

Class-imbalance-sensitive evaluation of automated decisions in mixed heart-lung auscultation

Lulu Wang
article en

Abstract

OBJECTIVE: Automated auscultation models are commonly evaluated using aggregate metrics, although mixed heart-lung recordings may be severely imbalanced and contain related acquisition conditions. This study examined whether prevalence, grouping, imbalance correction and decision-threshold selection materially affect the interpretation of cardiac and pulmonary abnormality decisions. Approach: A deterministic log-mel pipeline and nested grouped evaluation based on metadata-defined heart-lung sound-state combinations were applied to 145 mixed HLS-CMDS manikin recordings forming 60 operational groups. Independent and shared convolutional neural networks (CNNs) and a logistic-summary comparator were assessed using five repeated six-fold outer evaluations, with thresholds selected exclusively from inner grouped out-of-fold predictions. All-positive baselines, imbalance-training ablations, preprocessing sensitivity and support requirements for context analysis were evaluated. Results: The deterministic all-positive baselines achieved F1-scores of 0.953 for cardiac detection and 0.893 for pulmonary detection while failing to identify any normal recordings, demonstrating that high F1-score could be prevalence-driven. Across the five repeat-pooled outer evaluations, the logistic-summary comparator achieved the highest pulmonary F1-score, balanced accuracy and Matthews correlation coefficient (MCC), with mean ± SD values of 0.904 ± 0.015, 0.703 ± 0.041 and 0.451 ± 0.077, respectively. The independent CNN achieved the highest pulmonary specificity of 0.643 ± 0.143. Cardiac normal-class recovery remained unstable and did not support robust discrimination. For the independent pulmonary CNN, combined class weighting and oversampling did not improve balanced accuracy or MCC relative to class weighting alone. Support-aware analysis identified insufficient normal-context support for a reliable between-context comparison. Significance: Grouped partitioning, trivial baselines, nested threshold selection and normal-class-sensitive metrics materially changed performance interpretation. The results support methodological safeguards for automated auscultation evaluation but do not establish clinical diagnostic validity. .

Physiological Measurement
Reykjavík University (IS)
Reduced inequalities, Peace, Justice and strong institutions
Openalex Percentile: Top 11%
Phonocardiography and Auscultation Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.