Topology-enhanced machine learning for speech signal processing
In artificial-intelligence-aided signal processing, existing deep learning models often exhibit a black-box structure. Here, conceptually beyond spectral analysis, we demonstrate that topological methods not only effectively capture intrinsic and complex structural information but can also enhance neural networks. We provide a transparent methodology, TopCap, to capture topological features inherent in time series for basic machine learning. Compared to prior approaches, we obtain descriptors that probe finer information such as the vibration of a time series. Notably, in classifying voiced and voiceless consonants, TopCap achieves an accuracy consistently standing in comparison with neural network models. Moreover, by integrating TopCap features into those neural networks, our approach improves upon state-of-the-art methods in terms of robustness against noise, as well as accuracy, stability, convergence of loss function, and interpretability. Topological data analysis had shown promise in capturing intrinsic structural features of complex data. Here, the authors integrate topological data analysis with machine learning to capture structural features in consonant classification while enhancing interpretability and robustness against noise compared to traditional techniques.
Authors
- Yifei Zhu (ORCID: https://orcid.org/0000-0001-8918-1896)
- Zhiwang Yu (ORCID: https://orcid.org/0000-0003-3252-9067)
- Siheng Yi
- Haiyu Zhang (ORCID: https://orcid.org/0000-0002-4308-1918)
- Qingrui Qu
- Pingyao Feng
- Zeyang Ding
Institutions
- Southern University of Science and Technology (CN)
Publication Details
- Journal
- Nature Communications
- Published
- 2026-09-16
- DOI
- https://doi.org/10.1038/s41467-026-77649-z
- Primary Topic
- Topological and Geometric Data Analysis
- Type
- article
- Field-Weighted Citation Impact
- 0.00