The GPU Resilience Lottery: Understanding and Taming Real-World Hardware Errors in Heterogeneous AI Infrastructure

The booming LLM demand is driving a rapid increase in the scale and heterogeneity of modern AI infrastructure. To achieve high availability for sustained LLM workloads, infrastructure providers must contend with frequent and severe hardware errors. Without a holistic resilience understanding, providers may inadvertently schedule AI jobs onto GPUs with high error risks and mitigate errors with significant costs, further degrading cluster availability. As a result, AI jobs are forced to ‘win’ a GPU resilience lottery for interruption-free execution. However, betting on a winning ticket, i.e., selecting resilient GPU hardware and adopting efficient mitigations, requires a cross-layer understanding of error heterogeneity and severity, which remains opaque to current fault injection and fleet distribution studies. In this paper, we present a quantitative GPU resilience analysis in production clusters with 500+K GPUs, focusing on the multi-faceted heterogeneity and severity. We characterize distinct hardware errors via spatial patterns across distinct architectures and components, and temporal patterns correlated with deployment age. Furthermore, we quantify error severity across both hardware-layer outage, software-layer anomalies, and system-layer metrics. These insights lead to the design of LotteRig , a proactive availability enhancement framework that performs risk-aware GPU selection and cost-aware error mitigation. LotteRig achieves \(9.9\% \) fewer software interruptions and \(12.1\% \) higher hardware availability across diverse scenarios, substantially improving the odds of ‘winning’ the GPU resilience lottery.

Authors

Institutions

Publication Details

Journal
ACM Transactions on Architecture and Code Optimization
Published
2026-10-03
DOI
https://doi.org/10.1145/3849815
Primary Topic
Distributed systems and fault tolerance
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

The GPU Resilience Lottery: Understanding and Taming Real-World Hardware Errors in Heterogeneous AI Infrastructure

Lieven Eeckhout, Minyi Guo, Guodong Yang, Chao Li et al.
ACM Transactions on Architecture and Code Optimization
Distributed systems and fault tolerance
article

The GPU Resilience Lottery: Understanding and Taming Real-World Hardware Errors in Heterogeneous AI Infrastructure

Lieven Eeckhout, Minyi Guo, Guodong Yang, Chao Li, Liping Zhang, Xinkai Wang, Cheng Huang, Luping Wang, Yunwei Li
article en

Abstract

The booming LLM demand is driving a rapid increase in the scale and heterogeneity of modern AI infrastructure. To achieve high availability for sustained LLM workloads, infrastructure providers must contend with frequent and severe hardware errors. Without a holistic resilience understanding, providers may inadvertently schedule AI jobs onto GPUs with high error risks and mitigate errors with significant costs, further degrading cluster availability. As a result, AI jobs are forced to ‘win’ a GPU resilience lottery for interruption-free execution. However, betting on a winning ticket, i.e., selecting resilient GPU hardware and adopting efficient mitigations, requires a cross-layer understanding of error heterogeneity and severity, which remains opaque to current fault injection and fleet distribution studies. In this paper, we present a quantitative GPU resilience analysis in production clusters with 500+K GPUs, focusing on the multi-faceted heterogeneity and severity. We characterize distinct hardware errors via spatial patterns across distinct architectures and components, and temporal patterns correlated with deployment age. Furthermore, we quantify error severity across both hardware-layer outage, software-layer anomalies, and system-layer metrics. These insights lead to the design of LotteRig , a proactive availability enhancement framework that performs risk-aware GPU selection and cost-aware error mitigation. LotteRig achieves \(9.9\% \) fewer software interruptions and \(12.1\% \) higher hardware availability across diverse scenarios, substantially improving the odds of ‘winning’ the GPU resilience lottery.

ACM Transactions on Architecture and Code Optimization
Guizhou University (CN), Shanghai Jiao Tong University (CN), Ghent University (BE), Alibaba Group (China) (CN)
Openalex Percentile: Top 9%
Distributed systems and fault tolerance
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.