Full-memory storage data fault tolerance mechanism and rapid failure recovery system based on deep learning and neural networks
The rapid advancement of intelligent computing and data-intensive applications has accelerated the adoption of full-memory storage architectures which deliver fast data access and rapid data processing operations. The system suffers from high vulnerability because its hardware components and software components and system nodes can experience various types of faults. To address these challenges, this research proposes a Deep Learning (DL) and Neural Network–based Fault prediction and tolerance mechanism for full-memory environments. The system employs a Saplings Growing Up–tuned Residual Feedforward Neural Network (SGU-ResFNN), in which the Residual Feedforward Neural Network (ResFNN) leverages residual learning to improve gradient propagation, accelerate convergence, and reduce overfitting during fault prediction. The Saplings Growing Up (SGU) optimization algorithm adaptively fine-tunes network parameters, including learning rate, weight adjustments, and activation thresholds, to enhance model stability and precision under dynamic memory conditions.Classification label is 0 No fault, 1 Fault occurred. Min–Max normalized to standardize feature scales and eliminate bias. A Wavelet Transform (WT) based feature extraction mechanism decomposes fault signals into multi-resolution time–frequency representations, allowing accurate identification of transient and persistent fault patterns. Implemented in Python, The proposed system integrates Min–Max normalization, Wavelet Transform-based feature extraction, and SGU-ResFNN for fault prediction, evaluated on a 1000-sample dataset, achieving 95% accuracy, 92.6% recall, and 172 kWh power consumption. These results highlight the model’s effectiveness in enhancing fault tolerance and rapid recovery in full-memory storage environments. SGU-ResFNN is an improvement of fault prediction in full-memory storage systems based on deep learning. WT is a technique used to retrieve multiple-resolution faults to obtain faulty detection. 95% accuracy and low power consumption are achieved in high performance computing.
Authors
- Yuxuan Li (ORCID: https://orcid.org/0000-0003-0611-0764)
- Wu Hong (ORCID: https://orcid.org/0000-0001-9980-2481)
Institutions
- Jiangxi University of Water Resources and Electric Power (CN)
Publication Details
- Journal
- Discover Internet of Things
- Published
- 2026-09-05
- DOI
- https://doi.org/10.1007/s43926-026-00442-3
- Primary Topic
- Advanced Data Storage Technologies
- Type
- article
- Field-Weighted Citation Impact
- 0.00