Bounding the Human Interface in Clinical AI Safety Arguments: What Verification Cannot See
Clinical artificial intelligence has concentrated its safety effort on making the model trustworthy. Industrial safety engineering solved an equivalent problem differently: it does not require the unreliable component to become reliable, it architects the system so that the component's failure cannot propagate. We report the verification and validation of a clinical decision kernel built on that principle — a deterministic kernel whose governing rules contain no probabilistic step, paired with a probabilistic shell that holds no execution authority. Across 94,036 adversarial trials generated by four model families from three vendors (Google, Anthropic, OpenAI, including one reasoning-class control arm), attacking models produced 21,123 distinct clinical labels absent from the governed set. None altered a clinical decision once the architectural bound was in force. Attack coverage differed by a factor of 3.69 across vendors but only 1.27 within a vendor across model classes, indicating that coverage is determined by model family rather than model capability class — a methodological result with direct consequences for how adversarial testing of clinical systems should be designed. Verification succeeded. Validation did not return the same answer. Of 359 label instances that did alter a decision before the bound was applied, 346 (96.4%) were legitimate clinical writings that the system failed to recognise — lowercase, full-width, and zero-width-space variants of a governed term, together with simplified-Chinese forms and a real clinical abbreviation the systemhad no entry for. Under 2% were genuinely synthetic strings. The residual risk had not been eliminated; it had migrated from the AI-to-kernel interface to the human-machine interface, which is where systems engineering practice has long predicted that the dominant residual risk resides. We argue that hallucination containment in clinical AI is a tractable engineering problem, that it is tractable by verification rather than by model improvement, and that bounding one interface should be expected to relocate risk rather than remove it.
Authors
- Lu-An Chiu (ORCID: https://orcid.org/0009-0003-2165-0950)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-25
- DOI
- https://doi.org/10.5281/zenodo.22960590
- Primary Topic
- Adversarial Robustness in Machine Learning
- Type
- preprint