Beyond Endpoint Agreement: Trajectory Integrity, Criterion Recovery, and Retrospective Self-Audit in Nine Language-Model Systems Under Simulated ICU Triage

Large language models (LLMs) are often evaluated from isolated responses or final answers, although high-consequence use unfolds across interaction trajectories. We formalized trajectory integrity as a non-scalar profile separating endpoint recovery, criterion recovery, factual-state recovery, and retrospective fidelity, and applied it to nine LLM systems completing a simulated allocation of one intensive care unit (ICU) bed between two patients. Nine systems were evaluated in Spanish, English, German, and Simplified Chinese under four conditions, yielding 16 trajectories per system and 144 nested trajectories. Each trajectory elicited an allocation, governing criterion, prospective reversal threshold, responses to contextual pressures, an explicit reset, and retrospective audits. The model was the primary cross-model unit. Endpoint recovery after explicit reset occurred in 144/144 trajectories, but every system had at least one adjudicated trajectory with non-integral criterion recovery. Full factual recovery occurred in 132/144 trajectories; 12 were condition-specific, partial, or ambiguous. No system produced a fully faithful structured retrospective checksum across all 16 trajectories. Forced retrospective audits were completed in 144/144 trajectories, reflecting format compliance rather than spontaneous contradiction. A proposed composite metric failed semantic and arithmetic comparability across systems, supporting use of a profile rather than a single score. These results show that endpoint agreement can coexist with differences in recovered criterion, factual state, and retrospective account. The analysis is descriptive: each experimental cell contains one generation, the scenario is simulated, and several interpretive variables lack complete independent double-coding.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-29
DOI
https://doi.org/10.5281/zenodo.23045592
Primary Topic
Simulation-Based Education in Healthcare
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Beyond Endpoint Agreement: Trajectory Integrity, Criterion Recovery, and Retrospective Self-Audit in Nine Language-Model Systems Under Simulated ICU Triage

Evans Tovar
Zenodo (CERN European Organization for Nuclear Research)
Simulation-Based Education in Healthcare
preprint

Beyond Endpoint Agreement: Trajectory Integrity, Criterion Recovery, and Retrospective Self-Audit in Nine Language-Model Systems Under Simulated ICU Triage

Evans Tovar
preprint en

Abstract

Large language models (LLMs) are often evaluated from isolated responses or final answers, although high-consequence use unfolds across interaction trajectories. We formalized trajectory integrity as a non-scalar profile separating endpoint recovery, criterion recovery, factual-state recovery, and retrospective fidelity, and applied it to nine LLM systems completing a simulated allocation of one intensive care unit (ICU) bed between two patients. Nine systems were evaluated in Spanish, English, German, and Simplified Chinese under four conditions, yielding 16 trajectories per system and 144 nested trajectories. Each trajectory elicited an allocation, governing criterion, prospective reversal threshold, responses to contextual pressures, an explicit reset, and retrospective audits. The model was the primary cross-model unit. Endpoint recovery after explicit reset occurred in 144/144 trajectories, but every system had at least one adjudicated trajectory with non-integral criterion recovery. Full factual recovery occurred in 132/144 trajectories; 12 were condition-specific, partial, or ambiguous. No system produced a fully faithful structured retrospective checksum across all 16 trajectories. Forced retrospective audits were completed in 144/144 trajectories, reflecting format compliance rather than spontaneous contradiction. A proposed composite metric failed semantic and arithmetic comparability across systems, supporting use of a profile rather than a single score. These results show that endpoint agreement can coexist with differences in recovered criterion, factual state, and retrospective account. The analysis is descriptive: each experimental cell contains one generation, the scenario is simulated, and several interpretive variables lack complete independent double-coding.

Zenodo (CERN European Organization for Nuclear Research)
Peace, Justice and strong institutions
Simulation-Based Education in Healthcare
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Beyond Endpoint Agreement: Trajectory Integrity, Criterion Recovery, and Retrospective Self-Audit in Nine Language-Model Systems Under Simulated ICU Triage — Evans Tovar · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS