A Cohort-Level Floor with Limits to Fixed-Detector Transfer: A Pre-Registered Study of Calibrated Early-Response Geometry for Input-Label Discrimination Across Ten Language Models, with a Registered Six-Task Extension

We test whether early-response internal signals discriminate supplied input labels across language models, and which calibration choices transfer. Every prompt contains a candidate answer, hypothesis, summary or response; the target is its correctness, entailment or faithfulness, not whether the tested model hallucinates in free generation. We combine attention morphology, residual-stream motion, readout geometry and confidence in a 29-signal panel. A nested out-of-bag selector chooses a signal and its sign within each resample. The confirmatory protocols were pre-registered before data collection. On the two-task core across ten models, the confidence-free selector clears its registered discrimination criterion on 18/20 deployments (PASS, bar >=17/20); the full panel also scores 18/20 but misses its stricter >=19/20 bar. Twelve different geometric signals are selected. With component signs fixed from earlier, disjoint model-specific calibration, a fusion aggregate exceeds 0.55 AUROC on 9/10 ANLI and 10/10 TriviaQA model holdouts. This holds out current pooling data, not the model's entire calibration history. In the separately registered six-task extension, HaluEval-QA calibration passes 10/10 (bar >=8/10), while a detector with common component signs and a pooled outer sign passes 6/10 (same bar). Four holdouts reverse the observed signal direction. Both replication endpoints miss their >=17/20 bar at 7/20, following behavioral exclusions and a registered abort rule. These results support cohort-level discrimination under calibration and identify limits to the tested fixed detector; they do not exclude other transferable detectors. Post-registration operating-point analyses further limit practical interpretation: at a target 10% false-alarm rate (realized task medians 10-12%), median detection rates range from 0.25 to 0.70. Projected precision at 10% prevalence is 0.21-0.42. Clearing the registered criterion does not establish a usable alarm. Not peer reviewed.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-12
DOI
https://doi.org/10.5281/zenodo.22003458
Primary Topic
Neurobiology of Language and Bilingualism
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

A Cohort-Level Floor with Limits to Fixed-Detector Transfer: A Pre-Registered Study of Calibrated Early-Response Geometry for Input-Label Discrimination Across Ten Language Models, with a Registered Six-Task Extension

Michael S. R. Kitti
Zenodo (CERN European Organization for Nuclear Research)
Neurobiology of Language and Bilingualism
preprint

A Cohort-Level Floor with Limits to Fixed-Detector Transfer: A Pre-Registered Study of Calibrated Early-Response Geometry for Input-Label Discrimination Across Ten Language Models, with a Registered Six-Task Extension

Michael S. R. Kitti
preprint en

Abstract

We test whether early-response internal signals discriminate supplied input labels across language models, and which calibration choices transfer. Every prompt contains a candidate answer, hypothesis, summary or response; the target is its correctness, entailment or faithfulness, not whether the tested model hallucinates in free generation. We combine attention morphology, residual-stream motion, readout geometry and confidence in a 29-signal panel. A nested out-of-bag selector chooses a signal and its sign within each resample. The confirmatory protocols were pre-registered before data collection. On the two-task core across ten models, the confidence-free selector clears its registered discrimination criterion on 18/20 deployments (PASS, bar >=17/20); the full panel also scores 18/20 but misses its stricter >=19/20 bar. Twelve different geometric signals are selected. With component signs fixed from earlier, disjoint model-specific calibration, a fusion aggregate exceeds 0.55 AUROC on 9/10 ANLI and 10/10 TriviaQA model holdouts. This holds out current pooling data, not the model's entire calibration history. In the separately registered six-task extension, HaluEval-QA calibration passes 10/10 (bar >=8/10), while a detector with common component signs and a pooled outer sign passes 6/10 (same bar). Four holdouts reverse the observed signal direction. Both replication endpoints miss their >=17/20 bar at 7/20, following behavioral exclusions and a registered abort rule. These results support cohort-level discrimination under calibration and identify limits to the tested fixed detector; they do not exclude other transferable detectors. Post-registration operating-point analyses further limit practical interpretation: at a target 10% false-alarm rate (realized task medians 10-12%), median detection rates range from 0.25 to 0.70. Projected precision at 10% prevalence is 0.21-0.42. Clearing the registered criterion does not establish a usable alarm. Not peer reviewed.

Zenodo (CERN European Organization for Nuclear Research)
Independent Dance (GB), Oldham Council (GB)
Reduced inequalities
Neurobiology of Language and Bilingualism
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.