Bounding the Human Interface in Clinical AI Safety Arguments: What Verification Cannot See

Clinical artificial intelligence has concentrated its safety effort on making the model trustworthy. Industrial safety engineering solved an equivalent problem differently: it does not require the unreliable component to become reliable, it architects the system so that the component's failure cannot propagate. We report the verification and validation of a clinical decision kernel built on that principle — a deterministic kernel whose governing rules contain no probabilistic step, paired with a probabilistic shell that holds no execution authority. Across 94,036 adversarial trials generated by four model families from three vendors (Google, Anthropic, OpenAI, including one reasoning-class control arm), attacking models produced 21,123 distinct clinical labels absent from the governed set. None altered a clinical decision once the architectural bound was in force. Attack coverage differed by a factor of 3.69 across vendors but only 1.27 within a vendor across model classes, indicating that coverage is determined by model family rather than model capability class — a methodological result with direct consequences for how adversarial testing of clinical systems should be designed. Verification succeeded. Validation did not return the same answer. Of 359 label instances that did alter a decision before the bound was applied, 346 (96.4%) were legitimate clinical writings that the system failed to recognise — lowercase, full-width, and zero-width-space variants of a governed term, together with simplified-Chinese forms and a real clinical abbreviation the systemhad no entry for. Under 2% were genuinely synthetic strings. The residual risk had not been eliminated; it had migrated from the AI-to-kernel interface to the human-machine interface, which is where systems engineering practice has long predicted that the dominant residual risk resides. We argue that hallucination containment in clinical AI is a tractable engineering problem, that it is tractable by verification rather than by model improvement, and that bounding one interface should be expected to relocate risk rather than remove it.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-25
DOI
https://doi.org/10.5281/zenodo.22960589
Primary Topic
Adversarial Robustness in Machine Learning
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Bounding the Human Interface in Clinical AI Safety Arguments: What Verification Cannot See

Lu-An Chiu
Zenodo (CERN European Organization for Nuclear Research)
Adversarial Robustness in Machine Learning
preprint

Bounding the Human Interface in Clinical AI Safety Arguments: What Verification Cannot See

Lu-An Chiu
preprint en

Abstract

Clinical artificial intelligence has concentrated its safety effort on making the model trustworthy. Industrial safety engineering solved an equivalent problem differently: it does not require the unreliable component to become reliable, it architects the system so that the component's failure cannot propagate. We report the verification and validation of a clinical decision kernel built on that principle — a deterministic kernel whose governing rules contain no probabilistic step, paired with a probabilistic shell that holds no execution authority. Across 94,036 adversarial trials generated by four model families from three vendors (Google, Anthropic, OpenAI, including one reasoning-class control arm), attacking models produced 21,123 distinct clinical labels absent from the governed set. None altered a clinical decision once the architectural bound was in force. Attack coverage differed by a factor of 3.69 across vendors but only 1.27 within a vendor across model classes, indicating that coverage is determined by model family rather than model capability class — a methodological result with direct consequences for how adversarial testing of clinical systems should be designed. Verification succeeded. Validation did not return the same answer. Of 359 label instances that did alter a decision before the bound was applied, 346 (96.4%) were legitimate clinical writings that the system failed to recognise — lowercase, full-width, and zero-width-space variants of a governed term, together with simplified-Chinese forms and a real clinical abbreviation the systemhad no entry for. Under 2% were genuinely synthetic strings. The residual risk had not been eliminated; it had migrated from the AI-to-kernel interface to the human-machine interface, which is where systems engineering practice has long predicted that the dominant residual risk resides. We argue that hallucination containment in clinical AI is a tractable engineering problem, that it is tractable by verification rather than by model improvement, and that bounding one interface should be expected to relocate risk rather than remove it.

Zenodo (CERN European Organization for Nuclear Research)
Adversarial Robustness in Machine Learning
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Bounding the Human Interface in Clinical AI Safety Arguments: What Verification Cannot See — Lu-An Chiu · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS