One Percent of the Bank: Adaptive Item Selection Reads Targeted Contamination Up While the Whole Bank Barely Moves
Adaptive testing is entering language-model evaluation as an efficiency device, and its low item exposure is offered as incidental contamination protection (Li et al., 2026b). That argument assumes untargeted leakage, yet a maximum-information test asks predictable items that anyone observing administrations can learn. Here we fine-tune a 1.5B-parameter model on 1% of a self-calibrated MMLU bank, chosen by exposure in a clean adaptive run, and score adaptive, random and whole-bank protocols on the same response vectors. A 41-item adaptive test then rises by 1.47 to 1.59 logits over its own clean reading (1.63 to 1.75 above the clean whole-bank reading) while the whole bank reads slightly lower. Because this concentrated test also moves under untargeted fine-tuning drift, we compare it with a random-item set trained to the same loss of clean-item accuracy: on this model the adaptive reading exceeds that control's by +1.19 logits (95% CI +0.85 to +1.54), positive in all ten seed pairs. Every item of the clean test lies inside the contaminated set, and an exposure set rebuilt only from simulated administrations reproduces the shift. Two further models show the same sign pattern; the tested selection-side defences lower the bias only at a clean-RMSE cost. Adaptive-test logs therefore need a threat model for their observers, and adaptive scores a whole-bank reading beside them.
Authors
- Ya-Fen Yeh
- Guan-Yuan Chen (ORCID: https://orcid.org/0000-0003-3298-0624)
Institutions
- National Tsing Hua University (TW)
- North Carolina Exploring Cultural Heritage Online (US)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-25
- DOI
- https://doi.org/10.5281/zenodo.22959945
- Primary Topic
- Psychometric Methodologies and Testing
- Type
- preprint