Domain-specific unsupervised pre-training for robust respiratory sound classification

Abstract Automated auscultation using wearable devices is essential for remote respiratory monitoring, but deep learning models often struggle to generalize due to the severe scarcity of annotated abnormal respiratory sounds. Since directly using synthetic data for supervised training risks learning artifacts instead of true pathological features, we propose a robust three-phase pipeline leveraging massive synthetic data without synthetic labels. First, a modified StyleGAN2 natively synthesizes rectangular Mel-spectrograms to preserve high-temporal-resolution acoustic characteristics, validated by kernel audio distance. Second, 100,000 synthetic spectrograms are used for unsupervised variational autoencoder pre-training, introducing a parallel asymmetric convolutional block to independently capture distinct time and frequency semantics. Finally, the encoder is repurposed as a feature extractor to train a lightweight classifier on limited clinical data. Empirical evaluations demonstrate that while a serial asymmetric kernel geometry yields the highest classification accuracy under frequency-dominant pathologies, our parallel architecture achieves highly competitive performance using only 43% of conventional square baseline parameters and rivals the massive pretrained audio neural networks CNN14 model using merely 0.5% of its convolutional footprint. This framework provides a potential scalable pathway for domains with critically limited annotated data.

Authors

Institutions

Publication Details

Journal
Artificial Life and Robotics
Published
2026-09-08
DOI
https://doi.org/10.1007/s10015-026-01151-4
Primary Topic
Phonocardiography and Auscultation Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Domain-specific unsupervised pre-training for robust respiratory sound classification

Satoshi Konno, Yasumasa Tamura, Kaoruko Shimizu, Takehiro Hirasawa et al.
Artificial Life and Robotics
Phonocardiography and Auscultation Techniques
article

Domain-specific unsupervised pre-training for robust respiratory sound classification

Satoshi Konno, Yasumasa Tamura, Kaoruko Shimizu, Takehiro Hirasawa, Masahito Yamamoto
article en

Abstract

Abstract Automated auscultation using wearable devices is essential for remote respiratory monitoring, but deep learning models often struggle to generalize due to the severe scarcity of annotated abnormal respiratory sounds. Since directly using synthetic data for supervised training risks learning artifacts instead of true pathological features, we propose a robust three-phase pipeline leveraging massive synthetic data without synthetic labels. First, a modified StyleGAN2 natively synthesizes rectangular Mel-spectrograms to preserve high-temporal-resolution acoustic characteristics, validated by kernel audio distance. Second, 100,000 synthetic spectrograms are used for unsupervised variational autoencoder pre-training, introducing a parallel asymmetric convolutional block to independently capture distinct time and frequency semantics. Finally, the encoder is repurposed as a feature extractor to train a lightweight classifier on limited clinical data. Empirical evaluations demonstrate that while a serial asymmetric kernel geometry yields the highest classification accuracy under frequency-dominant pathologies, our parallel architecture achieves highly competitive performance using only 43% of conventional square baseline parameters and rivals the massive pretrained audio neural networks CNN14 model using merely 0.5% of its convolutional footprint. This framework provides a potential scalable pathway for domains with critically limited annotated data.

Artificial Life and Robotics
Hokkaido University (JP)
Openalex Percentile: Top 12%
Phonocardiography and Auscultation Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Domain-specific unsupervised pre-training for robust respiratory sound classification — Satoshi Konno, Yasumasa Tamura, et al. · Artificial Life and Robotics (2026) | TGRS Research Map | TGRS