Domain-Adaptive Audio Large Language Model for Acoustic Fault Diagnosis and Semantic Description of Coal Mine Equipment

Underground coal-mine equipment operates under broadband noise, high dust, humidity, and methane. Acoustic sensing is uniquely suited to this environment: it captures vibration, friction, and airflow signatures without physical contact, incurs low sensor-deployment cost, responds at millisecond speed, and remains effective in low-light, high-dust conditions where optical and vibration alternatives fail. Acoustic fault perception is therefore critically important in underground coal-mine operations. Three unresolved challenges remain: (i) motor whine, material-collision impacts, and ventilation-fan roar compound into a low-SNR soundscape where conventional models lose noise robustness; (ii) acoustic signatures vary widely across equipment types and fault-development stages; and (iii) existing supervised classifiers, trained on imbalanced data, exhibit limited generalization and output only binary judgments, lacking the semantic descriptions that maintenance crews actually need. To address all three, we propose a domain-adaptive audio LLM coupling a BEATs encoder (frozen during Stage II, adapted via Low-Rank Adaptation (LoRA) during Stage I), a Querying Transformer (Q-Former) alignment layer with Dynamic Acoustic Token Compression (DATC), a LLaMA-3.1-8B decoder adapted via LoRA, and Constrained Decoding for Structured Fault Description (CD-SFD) enforcing a three-slot output of fault type, danger level, and handling recommendation. DATC allocates query budget by signal energy to suppress noise-dominated frames; CD-SFD is a state-machine decoder that guarantees the three-slot schema. We release CMEASD: 1200 recordings comprising 24 physical machines (4 per equipment type, 12/6/6 machine-disjoint split). Under machine-disjoint evaluation, the model reaches Macro Accuracy 84.6 ± 1.5% and Macro F1 82.9 ± 1.6%, outperforming the strongest discriminative baseline (PANNs-Transformer, 81.7%) by +2.9 pp and SALMONN-LoRA by +2.5 pp. Ablations attribute +2.4/+1.4/+0.9 pp to DATC, CD-SFD, and the domain prompt. A 30-day mine trial achieves 7.0 s end-to-end latency with 21/30 days of stable, zero-false-shutdown operation.

Authors

Institutions

Publication Details

Journal
Algorithms
Published
2026-09-09
DOI
https://doi.org/10.3390/a19090778
Primary Topic
Machine Fault Diagnosis Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Domain-Adaptive Audio Large Language Model for Acoustic Fault Diagnosis and Semantic Description of Coal Mine Equipment

Daming Cui, Xin Zhang, Qiang Ma
Algorithms
Machine Fault Diagnosis Techniques
article

Domain-Adaptive Audio Large Language Model for Acoustic Fault Diagnosis and Semantic Description of Coal Mine Equipment

Daming Cui, Xin Zhang, Qiang Ma
article en

Abstract

Underground coal-mine equipment operates under broadband noise, high dust, humidity, and methane. Acoustic sensing is uniquely suited to this environment: it captures vibration, friction, and airflow signatures without physical contact, incurs low sensor-deployment cost, responds at millisecond speed, and remains effective in low-light, high-dust conditions where optical and vibration alternatives fail. Acoustic fault perception is therefore critically important in underground coal-mine operations. Three unresolved challenges remain: (i) motor whine, material-collision impacts, and ventilation-fan roar compound into a low-SNR soundscape where conventional models lose noise robustness; (ii) acoustic signatures vary widely across equipment types and fault-development stages; and (iii) existing supervised classifiers, trained on imbalanced data, exhibit limited generalization and output only binary judgments, lacking the semantic descriptions that maintenance crews actually need. To address all three, we propose a domain-adaptive audio LLM coupling a BEATs encoder (frozen during Stage II, adapted via Low-Rank Adaptation (LoRA) during Stage I), a Querying Transformer (Q-Former) alignment layer with Dynamic Acoustic Token Compression (DATC), a LLaMA-3.1-8B decoder adapted via LoRA, and Constrained Decoding for Structured Fault Description (CD-SFD) enforcing a three-slot output of fault type, danger level, and handling recommendation. DATC allocates query budget by signal energy to suppress noise-dominated frames; CD-SFD is a state-machine decoder that guarantees the three-slot schema. We release CMEASD: 1200 recordings comprising 24 physical machines (4 per equipment type, 12/6/6 machine-disjoint split). Under machine-disjoint evaluation, the model reaches Macro Accuracy 84.6 ± 1.5% and Macro F1 82.9 ± 1.6%, outperforming the strongest discriminative baseline (PANNs-Transformer, 81.7%) by +2.9 pp and SALMONN-LoRA by +2.5 pp. Ablations attribute +2.4/+1.4/+0.9 pp to DATC, CD-SFD, and the domain prompt. A 30-day mine trial achieves 7.0 s end-to-end latency with 21/30 days of stable, zero-false-shutdown operation.

AlgorithmsVol. 19(9)
Shaanxi Coal Chemical Industry Technology Research Institute (CN)
Peace, Justice and strong institutions
Openalex Percentile: Top 24%
Machine Fault Diagnosis Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.