A Safe Stop Can Have the Wrong Diagnosis

Protective controls in AI systems are often evaluated as if a terminal label also explains why the system stopped. That assumption can erase an important distinction: a stop may be prudent while its causal diagnosis is wrong; a local harness failure may occur before the target is measured; and a validator may accept a representation that later direct evidence shows to be invalid. This paper applies five pre-frozen, forced-choice primary incident classes to seven de-identified naturalistic incidents while treating upstream causal mechanisms as potentially overlapping. Three model-generated coding returns, produced from the same frozen source-key-withheld packet, applied the primary classes. Six of seven cases received unanimous primary labels, and 19 of 21 within-case return pairs agreed (90.5%); nominal Krippendorff alpha was 0.88 and is reported descriptively because the frozen packet specified no adequacy threshold. The sole primary disagreement concerned CASE-02; under the pre-coding source key, stopping remained warranted after the unsupported diagnosis was removed, and the key assigned the case to SAFE_HOLD_WRONG_DIAGNOSIS. Secondary fields were less stable, especially prior-valid-substate preservation. The study is deliberately small and selected. Its narrower result is methodological: terminal stop prudence and causal-diagnosis support remained distinct, reviewable coding objects in this seven-case application exercise. The agreement statistics describe consistency across model-generated returns under one coding instrument; they do not estimate agreement among independent human coders.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-25
DOI
https://doi.org/10.5281/zenodo.22954804
Primary Topic
Explainable Artificial Intelligence (XAI)
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

A Safe Stop Can Have the Wrong Diagnosis

Logan Davis
Zenodo (CERN European Organization for Nuclear Research)
Explainable Artificial Intelligence (XAI)
preprint

A Safe Stop Can Have the Wrong Diagnosis

Logan Davis
preprint en

Abstract

Protective controls in AI systems are often evaluated as if a terminal label also explains why the system stopped. That assumption can erase an important distinction: a stop may be prudent while its causal diagnosis is wrong; a local harness failure may occur before the target is measured; and a validator may accept a representation that later direct evidence shows to be invalid. This paper applies five pre-frozen, forced-choice primary incident classes to seven de-identified naturalistic incidents while treating upstream causal mechanisms as potentially overlapping. Three model-generated coding returns, produced from the same frozen source-key-withheld packet, applied the primary classes. Six of seven cases received unanimous primary labels, and 19 of 21 within-case return pairs agreed (90.5%); nominal Krippendorff alpha was 0.88 and is reported descriptively because the frozen packet specified no adequacy threshold. The sole primary disagreement concerned CASE-02; under the pre-coding source key, stopping remained warranted after the unsupported diagnosis was removed, and the key assigned the case to SAFE_HOLD_WRONG_DIAGNOSIS. Secondary fields were less stable, especially prior-valid-substate preservation. The study is deliberately small and selected. Its narrower result is methodological: terminal stop prudence and causal-diagnosis support remained distinct, reviewable coding objects in this seven-case application exercise. The agreement statistics describe consistency across model-generated returns under one coding instrument; they do not estimate agreement among independent human coders.

Zenodo (CERN European Organization for Nuclear Research)
Peace, Justice and strong institutions
Explainable Artificial Intelligence (XAI)
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

A Safe Stop Can Have the Wrong Diagnosis — Logan Davis · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS