A Cohort-Level Floor with Limits to Fixed-Detector Transfer: A Pre-Registered Study of Calibrated Early-Response Geometry for Input-Label Discrimination Across Ten Language Models, with a Registered Six-Task Extension
We test whether early-response internal signals discriminate supplied input labels across language models, and which calibration choices transfer. Every prompt contains a candidate answer, hypothesis, summary or response; the target is its correctness, entailment or faithfulness, not whether the tested model hallucinates in free generation. We combine attention morphology, residual-stream motion, readout geometry and confidence in a 29-signal panel. A nested out-of-bag selector chooses a signal and its sign within each resample. The confirmatory protocols were pre-registered before data collection. On the two-task core across ten models, the confidence-free selector clears its registered discrimination criterion on 18/20 deployments (PASS, bar >=17/20); the full panel also scores 18/20 but misses its stricter >=19/20 bar. Twelve different geometric signals are selected. With component signs fixed from earlier, disjoint model-specific calibration, a fusion aggregate exceeds 0.55 AUROC on 9/10 ANLI and 10/10 TriviaQA model holdouts. This holds out current pooling data, not the model's entire calibration history. In the separately registered six-task extension, HaluEval-QA calibration passes 10/10 (bar >=8/10), while a detector with common component signs and a pooled outer sign passes 6/10 (same bar). Four holdouts reverse the observed signal direction. Both replication endpoints miss their >=17/20 bar at 7/20, following behavioral exclusions and a registered abort rule. These results support cohort-level discrimination under calibration and identify limits to the tested fixed detector; they do not exclude other transferable detectors. Post-registration operating-point analyses further limit practical interpretation: at a target 10% false-alarm rate (realized task medians 10-12%), median detection rates range from 0.25 to 0.70. Projected precision at 10% prevalence is 0.21-0.42. Clearing the registered criterion does not establish a usable alarm. Not peer reviewed.
Authors
- Michael S. R. Kitti
Institutions
- Independent Dance (GB)
- Oldham Council (GB)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-12
- DOI
- https://doi.org/10.5281/zenodo.22003458
- Primary Topic
- Neurobiology of Language and Bilingualism
- Type
- preprint