The GPU Resilience Lottery: Understanding and Taming Real-World Hardware Errors in Heterogeneous AI Infrastructure
The booming LLM demand is driving a rapid increase in the scale and heterogeneity of modern AI infrastructure. To achieve high availability for sustained LLM workloads, infrastructure providers must contend with frequent and severe hardware errors. Without a holistic resilience understanding, providers may inadvertently schedule AI jobs onto GPUs with high error risks and mitigate errors with significant costs, further degrading cluster availability. As a result, AI jobs are forced to ‘win’ a GPU resilience lottery for interruption-free execution. However, betting on a winning ticket, i.e., selecting resilient GPU hardware and adopting efficient mitigations, requires a cross-layer understanding of error heterogeneity and severity, which remains opaque to current fault injection and fleet distribution studies. In this paper, we present a quantitative GPU resilience analysis in production clusters with 500+K GPUs, focusing on the multi-faceted heterogeneity and severity. We characterize distinct hardware errors via spatial patterns across distinct architectures and components, and temporal patterns correlated with deployment age. Furthermore, we quantify error severity across both hardware-layer outage, software-layer anomalies, and system-layer metrics. These insights lead to the design of LotteRig , a proactive availability enhancement framework that performs risk-aware GPU selection and cost-aware error mitigation. LotteRig achieves \(9.9\% \) fewer software interruptions and \(12.1\% \) higher hardware availability across diverse scenarios, substantially improving the odds of ‘winning’ the GPU resilience lottery.
Authors
- Lieven Eeckhout (ORCID: https://orcid.org/0000-0001-8792-4473)
- Minyi Guo (ORCID: https://orcid.org/0000-0003-0034-2302)
- Guodong Yang (ORCID: https://orcid.org/0000-0003-1908-071X)
- Chao Li (ORCID: https://orcid.org/0000-0001-6218-4659)
- Liping Zhang (ORCID: https://orcid.org/0000-0003-2334-3471)
- Xinkai Wang (ORCID: https://orcid.org/0000-0003-3764-8065)
- Cheng Huang (ORCID: https://orcid.org/0009-0006-1313-4464)
- Luping Wang (ORCID: https://orcid.org/0009-0000-1187-0202)
- Yunwei Li (ORCID: https://orcid.org/0009-0007-4082-2307)
Institutions
- Guizhou University (CN)
- Shanghai Jiao Tong University (CN)
- Ghent University (BE)
- Alibaba Group (China) (CN)
Publication Details
- Journal
- ACM Transactions on Architecture and Code Optimization
- Published
- 2026-10-03
- DOI
- https://doi.org/10.1145/3849815
- Primary Topic
- Distributed systems and fault tolerance
- Type
- article
- Field-Weighted Citation Impact
- 0.00