When External Numerical Guards Detect Injected Faults in Neural-Network Inference
External guards can withhold an accelerator's output and trigger recovery, but their usefulness depends on which faults their numerical tests detect. We examine software fault-injection campaigns in convolutional networks, a vision transformer, and a decoder language model, distinguishing reported results from independently recounted records. Removing ReLU6 increased reported coverage from 0/155 to 325/326 harmful events at a tested MobileNetV2 site; adding clipping to a ResNet-50 site reduced it from 634/634 to 0/20. These interventions support a causal influence of activation restriction, while also changing which injected faults are harmful. In the BF16 Qwen3-4B RUN07c study, the recounted same-position coverage was 681/727. Recovered multi-boundary records contained 623 same-position harmful events: the first downstream boundary detected 576 and the union detected 577, differing from the smaller count in the supplied manuscript. Weight-only FP8 RUN09 records confirmed no first-boundary or union K20 detections among 288 finite harmful events, whereas BF16 detected 689/750. A separate paired RUN07 completed 780,000 trials using BF16 and per-tensor FP8 weights-and-activations storage with BF16 arithmetic. Its FP8 logit-summary detector detected 0/1,934 finite harmful events, while its internal residual-stream K20 detector did detect 231/1,866. This makes the limitation specific to the detector and numerical contract. We distinguish these observations from a universal observability rule: a bounded operation need not place faulty outputs inside a detector's acceptance region, and an unrestricted path need not make every harmful perturbation detectable. The results motivate graph-informed placement followed by empirical qualification under a declared fault sampler, numerical configuration, profile, and observation scope. They do not establish radiation-event rates, physical containment performance, or protection against persistent cache faults during incremental decoding. Remaining provenance and configuration checks are specified before release. Results are identified as reported or recounted; recounted values were verified against saved trial records in October 2026. Software fault injection alone does not establish radiation or flight readiness.
Authors
- Serhii Serhieiev
- Lidiia Pukhkan
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-05
- DOI
- https://doi.org/10.5281/zenodo.23146540
- Primary Topic
- Software Testing and Debugging Techniques
- Type
- preprint