Joint fault prediction and remaining useful life estimation for large-scale mining equipment based on temporal convolutional network and multi-head self-attention
Abstract Reliable operation of large-scale mining equipment governs both production efficiency and the safety of the crews working underground, yet the maintenance strategies still common in the sector either replace healthy parts on a calendar or wait for a breakdown that arrives without warning. This paper proposes a joint fault prediction and remaining useful life (RUL) estimation method that couples temporal convolutional networks (TCN) with multi-head self-attention (MHSA). A multi-scale temporal convolution module built from three parallel dilated convolution branches with heterogeneous kernel sizes and dilation configurations resolves high-frequency fault impulses, medium-range operational trends and long-horizon degradation drift at the same time, while a channel attention mechanism weights each scale according to the input. The resulting features enter an MHSA-based fusion encoder that models global temporal dependencies across the observation window, so that the shared representation is at once locally discriminative and globally coherent. A unified multi-task framework with uncertainty-based dynamic loss weighting lets the fault classification head and the RUL regression head co-train over that representation without manual balancing. Because run-to-failure records from operating mine hoists are not publicly available, the method is validated on two public bearing benchmarks that are widely used as surrogates for hoist and gearbox bearings: the Case Western Reserve University data set for classification and the XJTU-SY accelerated degradation data set for RUL, in which one cycle denotes one minute of operation. Averaged over five random seeds, the method attains 98.72 ± 0.15% classification accuracy and an RUL RMSE of 11.83 ± 0.36 cycles, ahead of standalone TCN, Transformer, CNN-LSTM and classical machine-learning baselines as well as of four re-implemented integrated models, at the smallest parameter count of the comparison. Ablation experiments show that multi-scale branching, global attention encoding and joint training each contribute, the multi-task mechanism reducing the asymmetric score by 19.5% relative to single-task training. Robustness is further examined under impulsive shocks, power-frequency interference and sensor drift, and the step from these laboratory benchmarks to an operating mine is discussed as the principal open question.
Authors
- Fuyong Yang
- Huiyi Zhu
- Wenjun Xu (ORCID: https://orcid.org/0000-0002-6556-1728)
- Jianhui Mao
- Dongfang Li
Institutions
- Shanghai University of Engineering Science (CN)
- Quzhou College of Technology (CN)
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-09-18
- DOI
- https://doi.org/10.1038/s41598-026-69929-x
- Primary Topic
- Machine Fault Diagnosis Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00