Silicon Senescence: Gompertz Mortality in a Fielded GPU Population
Modern datacenters hold thousands of costly accelerators and yet operators have little evidence for how failure risk changes with service age. We reanalyze 30,207 NVIDIA K20X GPUs from the Oak Ridge Titan supercomputer to ask whether the field data contains an aging signal and what it could mean operationally. In the 11,370 GPUs first installed in 2016 or later, the data favor a Gompertz hazard with b = 0.472 yr⁻¹, but the late-age rise is concentrated in one 485-device serial-number group. Removing that group eliminates the model preference which shows that the fleet-level aging curves can mix true wearout with serial-group and history. The practical consequence is therefore an operating question. If voltage, temperature, and accumulated wear raise failure risk, can a small performance concession make the computing hardware last longer? In a separate three-year simulation we find that a fixed rated operating point lowers aggregate work by only 0.75 percentage points relative to normal default boost setting while reducing GPU removals from 27.3 to 13.8 per 120 device slots; one sampled wear-aware setting delivers 101.40 percent of rated work while reducing removals to 15.7. These are simulated, not measured, life-extension results. The point is the scale of the possible trade: changes in compute of order one percent can accompany changes in hardware loss of order tens of percent. Further testing on that possibility now requires modern fleet data that record serial group, age, operating stress, and failure together.
Authors
- Bruce H. Dean (ORCID: https://orcid.org/0009-0008-8153-3269)
Institutions
- Symplectic (UK) (GB)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-16
- DOI
- https://doi.org/10.5281/zenodo.22801443
- Primary Topic
- Advanced Data Storage Technologies
- Type
- preprint