A Safe Stop Can Have the Wrong Diagnosis
Protective controls in AI systems are often evaluated as if a terminal label also explains why the system stopped. That assumption can erase an important distinction: a stop may be prudent while its causal diagnosis is wrong; a local harness failure may occur before the target is measured; and a validator may accept a representation that later direct evidence shows to be invalid. This paper applies five pre-frozen, forced-choice primary incident classes to seven de-identified naturalistic incidents while treating upstream causal mechanisms as potentially overlapping. Three model-generated coding returns, produced from the same frozen source-key-withheld packet, applied the primary classes. Six of seven cases received unanimous primary labels, and 19 of 21 within-case return pairs agreed (90.5%); nominal Krippendorff alpha was 0.88 and is reported descriptively because the frozen packet specified no adequacy threshold. The sole primary disagreement concerned CASE-02; under the pre-coding source key, stopping remained warranted after the unsupported diagnosis was removed, and the key assigned the case to SAFE_HOLD_WRONG_DIAGNOSIS. Secondary fields were less stable, especially prior-valid-substate preservation. The study is deliberately small and selected. Its narrower result is methodological: terminal stop prudence and causal-diagnosis support remained distinct, reviewable coding objects in this seven-case application exercise. The agreement statistics describe consistency across model-generated returns under one coding instrument; they do not estimate agreement among independent human coders.
Authors
- Logan Davis (ORCID: https://orcid.org/0009-0006-8244-5610)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-25
- DOI
- https://doi.org/10.5281/zenodo.22954804
- Primary Topic
- Explainable Artificial Intelligence (XAI)
- Type
- preprint