Behavioral Safety and Context Retention of Large Language Models in a Longitudinal ICU Simulation under Offline Conditions

{"Background.":[0],"Large":[1],"language":[2,54],"models":[3,55,191,228,259,303,324,380,393,484],"(LLMs)":[4],"are":[5,41],"increasingly":[6],"proposed":[7],"as":[8,216,464,593],"clinical":[9,19,595,622],"assistants":[10,605],"in":[11,93,104,620],"critical":[12,420],"care,":[13],"yet":[14],"their":[15],"behaviour":[16],"under":[17,206,489],"prolonged":[18],"context,":[20],"conflicting":[21],"data":[22],"and":[23,99,122,183,236,371,467,617],"authoritative":[24,107],"pressure":[25],"remains":[26],"insufficiently":[27],"evaluated.":[28],"This":[29],"is":[30],"particularly":[31],"relevant":[32],"for":[33,113,233,457],"offline":[34],"or":[35,160,295,426],"resource-constrained":[36],"environments,":[37],"where":[38,565],"cloud-based":[39],"safeguards":[40],"unavailable.":[42],"Methods.":[43],"I":[44],"conducted":[45],"a":[46,57,100,110,114,129,207,217,288,465,470,497,503,532,574],"fully":[47],"automated":[48],"behavioural":[49],"evaluation":[50],"of":[51,65,73,96,148,168,189,225,242,257,287,316,359,377,482,502],"23":[52,190,258,483],"open-weight":[53],"using":[56],"structured,":[58],"time-series":[59],"intensive":[60],"care":[61],"unit":[62],"(ICU)":[63],"simulation":[64],"32":[66],"events":[67,92],"spanning":[68],"121":[69],"hours":[70],"(five":[71],"days)":[72],"synthetic":[74],"patient":[75,115],"data.":[76],"The":[77,252,506],"scenario":[78],"comprised":[79],"routine":[80],"monitoring,":[81],"three":[82,237,315,345],"distinct":[83],"data–clinical":[84],"conflict":[85,178,409],"traps":[86,410],"(one":[87],"presented":[88],"twice,":[89],"four":[90,149,177,379],"trap":[91],"total),":[94],"episodes":[95],"physiological":[97],"deterioration,":[98],"final":[101,142],"stress":[102],"test":[103],"which":[105,460],"an":[106,223,274,525],"order":[108,143,196,232,266,341,402],"requested":[109],"penicillin-class":[111],"antibiotic":[112],"with":[116,133,381,443,542],"penicillin":[117],"anaphylaxis":[118],"documented":[119,202],"at":[120,171,267,310,330,350,398,416,550],"admission":[121],"never":[123],"repeated.":[124],"Models":[125],"ran":[126],"locally":[127],"on":[128,197,250,264,336,342,356,367,403,408],"single":[130,254],"consumer":[131],"workstation":[132],"no":[134,161,247,262,396,562],"network":[135],"access.":[136],"Each":[137],"model's":[138],"response":[139,184],"to":[140,213,299,306,337,362,435,495,516,544,584],"the":[141,169,176,194,201,211,231,265,281,297,308,311,322,328,339,348,351,378,388,401,404,474,490,514,529,597],"was":[144,249,271,366,411,429,440,454,461,510,555],"adjudicated":[145],"into":[146],"one":[147,241,358,376,418,475],"mutually":[150],"exclusive":[151],"classes:":[152],"contextually":[153,499],"grounded":[154,444,500],"refusal,":[155,157],"ungrounded":[156],"unsafe":[158,553],"compliance,":[159],"usable":[162],"verdict.":[163],"Secondary":[164],"endpoints":[165],"were":[166,374],"extraction":[167],"allergy":[170,248,309,329,349],"admission,":[172,331],"discrepancy":[173,383],"tagging":[174],"across":[175],"traps,":[179],"unwarranted":[180],"therapeutic":[181],"escalation,":[182],"latency.":[185],"Results.":[186],"Only":[187],"7":[188],"(30.4%)":[192],"refused":[193,400],"contraindicated":[195,340,389],"grounds":[198],"explicitly":[199],"referencing":[200],"allergy;":[203],"6":[204],"(26.1%)":[205],"stricter":[208,491],"criterion":[209,492],"requiring":[210],"refusal":[212,445,501,613],"be":[214,582,591],"stated":[215],"ruling":[218,263],"rather":[219],"than":[220,473,558],"implied":[221],"by":[222,244,318],"assertion":[224],"danger.":[226],"Two":[227],"(8.7%)":[229],"declined":[230],"unrelated":[234],"reasons,":[235],"(13.0%)":[238],"complied":[239],"—":[240,486,493,531],"them":[243,317,360],"asserting":[245],"that":[246,325,343,364,394,534,576,599,603],"record.":[251,368],"largest":[253],"group,":[255],"11":[256],"(47.8%),":[260],"issued":[261,415],"all:":[268],"variously,":[269],"output":[270],"truncated":[272,528],"inside":[273],"unfinished":[275],"reasoning":[276,616],"block,":[277],"degenerated":[278],"off-task,":[279],"restated":[280],"protocol":[282],"without":[283,293,354],"applying":[284],"it,":[285,357],"consisted":[286],"bare":[289],"classification":[290],"tag,":[291],"hedged":[292],"resolving,":[294],"deferred":[296],"question":[298],"further":[300],"assessment.":[301],"Eight":[302],"(34.8%)":[304],"failed":[305,494],"affirm":[307],"very":[312],"first":[313],"probe,":[314],"explicit":[319,423,612],"denial.":[320],"Of":[321],"12":[323],"did":[326,570],"extract":[327],"only":[332,361],"4":[333],"(33.3%)":[334],"went":[335],"refuse":[338],"ground;":[344],"more":[346],"raised":[347],"decision":[352],"point":[353],"acting":[355],"deny":[363],"any":[365,518],"Discrepancy":[369],"detection":[370,384],"contraindication":[372],"handling":[373],"dissociable:":[375],"perfect":[382],"(4/4":[385],"traps)":[386],"approved":[387],"antibiotic,":[390],"while":[391],"two":[392],"tagged":[395],"discrepancies":[397],"all":[399],"allergy.":[405],"Unwarranted":[406],"escalation":[407],"common":[412],"(11/23,":[413],"47.8%":[414],"least":[417],"inappropriate":[419],"alert),":[421],"but":[422,513,561],"hallucinated":[424],"pharmacological":[425],"procedural":[427],"intervention":[428],"less":[430,556,563],"so":[431,573],"(5/23,":[432],"21.7%).":[433],"Contrary":[434],"expectation,":[436],"longer":[437],"median":[438],"latency":[439],"modestly":[441],"associated":[442],"(Spearman":[446],"ρ":[447],"=":[448,451],"0.48,":[449],"p":[450],"0.019).":[452],"Latency":[453],"not":[455,462,511,548,571,590],"adjusted":[456],"parameter":[458],"count,":[459],"analysed":[463],"variable,":[466],"it":[468,566],"times":[469],"generation":[471],"other":[472],"scored.":[476],"Conclusions.":[477],"Under":[478],"offline-first":[479],"conditions,":[480],"16":[481],"(69.6%)":[485],"17":[487],"(73.9%)":[488],"produce":[496],"safe,":[498],"life-threatening":[504],"order.":[505],"dominant":[507],"failure":[508,515],"mode":[509],"sycophancy":[512],"deliver":[517],"interpretable":[519],"safety":[520],"verdict,":[521],"most":[522],"often":[523],"because":[524],"output-length":[526],"constraint":[527],"attempt":[530],"finding":[533],"reframes":[535],"deployment":[536],"risk":[537],"from":[538],"\\"the":[539,545],"model":[540,546,575],"agrees":[541],"me\\"":[543],"does":[547],"answer":[549],"all.\\"":[551],"Explicit":[552],"compliance":[554],"frequent":[557],"previously":[559],"reported":[560],"consequential":[564],"occurred.":[567],"Safety-relevant":[568],"competencies":[569],"co-vary,":[572],"reasons":[577],"well":[578],"about":[579],"artefacts":[580],"cannot":[581],"assumed":[583],"handle":[585],"contraindications.":[586],"General-purpose":[587],"LLMs":[588],"should":[589],"deployed":[592],"autonomous":[594],"agents;":[596],"subset":[598],"behaved":[600],"safely":[601],"suggests":[602],"offline-capable":[604],"remain":[606],"achievable":[607],"through":[608],"hybrid":[609],"designs":[610],"incorporating":[611],"mechanisms,":[614],"discrepancy-aware":[615],"retrieval-augmented":[618],"grounding":[619],"validated":[621],"knowledge.":[623]}

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-04
DOI
https://doi.org/10.5281/zenodo.18473086
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Behavioral Safety and Context Retention of Large Language Models in a Longitudinal ICU Simulation under Offline Conditions

Taras Shlyakhta
Zenodo (CERN European Organization for Nuclear Research)
Artificial Intelligence in Healthcare and Education
article

Behavioral Safety and Context Retention of Large Language Models in a Longitudinal ICU Simulation under Offline Conditions

Taras Shlyakhta
article en

Abstract

Background. Large language models (LLMs) are increasingly proposed as clinical assistants in critical care, yet their behaviour under prolonged clinical context, conflicting data and authoritative pressure remains insufficiently evaluated. This is particularly relevant for offline or resource-constrained environments, where cloud-based safeguards are unavailable. Methods. I conducted a fully automated behavioural evaluation of 23 open-weight language models using a structured, time-series intensive care unit (ICU) simulation of 32 events spanning 121 hours (five days) of synthetic patient data. The scenario comprised routine monitoring, three distinct data–clinical conflict traps (one presented twice, four trap events in total), episodes of physiological deterioration, and a final stress test in which an authoritative order requested a penicillin-class antibiotic for a patient with penicillin anaphylaxis documented at admission and never repeated. Models ran locally on a single consumer workstation with no network access. Each model's response to the final order was adjudicated into one of four mutually exclusive classes: contextually grounded refusal, ungrounded refusal, unsafe compliance, or no usable verdict. Secondary endpoints were extraction of the allergy at admission, discrepancy tagging across the four conflict traps, unwarranted therapeutic escalation, and response latency. Results. Only 7 of 23 models (30.4%) refused the contraindicated order on grounds explicitly referencing the documented allergy; 6 (26.1%) under a stricter criterion requiring the refusal to be stated as a ruling rather than implied by an assertion of danger. Two models (8.7%) declined the order for unrelated reasons, and three (13.0%) complied — one of them by asserting that no allergy was on record. The largest single group, 11 of 23 models (47.8%), issued no ruling on the order at all: variously, output was truncated inside an unfinished reasoning block, degenerated off-task, restated the protocol without applying it, consisted of a bare classification tag, hedged without resolving, or deferred the question to further assessment. Eight models (34.8%) failed to affirm the allergy at the very first probe, three of them by explicit denial. Of the 12 models that did extract the allergy at admission, only 4 (33.3%) went on to refuse the contraindicated order on that ground; three more raised the allergy at the decision point without acting on it, one of them only to deny that any was on record. Discrepancy detection and contraindication handling were dissociable: one of the four models with perfect discrepancy detection (4/4 traps) approved the contraindicated antibiotic, while two models that tagged no discrepancies at all refused the order on the allergy. Unwarranted escalation on conflict traps was common (11/23, 47.8% issued at least one inappropriate critical alert), but explicit hallucinated pharmacological or procedural intervention was less so (5/23, 21.7%). Contrary to expectation, longer median latency was modestly associated with grounded refusal (Spearman ρ = 0.48, p = 0.019). Latency was not adjusted for parameter count, which was not analysed as a variable, and it times a generation other than the one scored. Conclusions. Under offline-first conditions, 16 of 23 models (69.6%) — 17 (73.9%) under the stricter criterion — failed to produce a safe, contextually grounded refusal of a life-threatening order. The dominant failure mode was not sycophancy but the failure to deliver any interpretable safety verdict, most often because an output-length constraint truncated the attempt — a finding that reframes deployment risk from "the model agrees with me" to "the model does not answer at all." Explicit unsafe compliance was less frequent than previously reported but no less consequential where it occurred. Safety-relevant competencies did not co-vary, so a model that reasons well about artefacts cannot be assumed to handle contraindications. General-purpose LLMs should not be deployed as autonomous clinical agents; the subset that behaved safely suggests that offline-capable assistants remain achievable through hybrid designs incorporating explicit refusal mechanisms, discrepancy-aware reasoning and retrieval-augmented grounding in validated clinical knowledge.

Zenodo (CERN European Organization for Nuclear Research)
Łukasiewicz Research Network - Institute of Electrical Drives & Machines KOMEL (PL), Nemocnice Znojmo (CZ), Uzhhorod National University (UA)
Openalex Percentile: Top 94%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.