One Percent of the Bank: Adaptive Item Selection Reads Targeted Contamination Up While the Whole Bank Barely Moves

Adaptive testing is entering language-model evaluation as an efficiency device, and its low item exposure is offered as incidental contamination protection (Li et al., 2026b). That argument assumes untargeted leakage, yet a maximum-information test asks predictable items that anyone observing administrations can learn. Here we fine-tune a 1.5B-parameter model on 1% of a self-calibrated MMLU bank, chosen by exposure in a clean adaptive run, and score adaptive, random and whole-bank protocols on the same response vectors. A 41-item adaptive test then rises by 1.47 to 1.59 logits over its own clean reading (1.63 to 1.75 above the clean whole-bank reading) while the whole bank reads slightly lower. Because this concentrated test also moves under untargeted fine-tuning drift, we compare it with a random-item set trained to the same loss of clean-item accuracy: on this model the adaptive reading exceeds that control's by +1.19 logits (95% CI +0.85 to +1.54), positive in all ten seed pairs. Every item of the clean test lies inside the contaminated set, and an exposure set rebuilt only from simulated administrations reproduces the shift. Two further models show the same sign pattern; the tested selection-side defences lower the bias only at a clean-RMSE cost. Adaptive-test logs therefore need a threat model for their observers, and adaptive scores a whole-bank reading beside them.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-25
DOI
https://doi.org/10.5281/zenodo.22959945
Primary Topic
Psychometric Methodologies and Testing
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

One Percent of the Bank: Adaptive Item Selection Reads Targeted Contamination Up While the Whole Bank Barely Moves

Ya-Fen Yeh, Guan-Yuan Chen
Zenodo (CERN European Organization for Nuclear Research)
Psychometric Methodologies and Testing
preprint

One Percent of the Bank: Adaptive Item Selection Reads Targeted Contamination Up While the Whole Bank Barely Moves

Ya-Fen Yeh, Guan-Yuan Chen
preprint en

Abstract

Adaptive testing is entering language-model evaluation as an efficiency device, and its low item exposure is offered as incidental contamination protection (Li et al., 2026b). That argument assumes untargeted leakage, yet a maximum-information test asks predictable items that anyone observing administrations can learn. Here we fine-tune a 1.5B-parameter model on 1% of a self-calibrated MMLU bank, chosen by exposure in a clean adaptive run, and score adaptive, random and whole-bank protocols on the same response vectors. A 41-item adaptive test then rises by 1.47 to 1.59 logits over its own clean reading (1.63 to 1.75 above the clean whole-bank reading) while the whole bank reads slightly lower. Because this concentrated test also moves under untargeted fine-tuning drift, we compare it with a random-item set trained to the same loss of clean-item accuracy: on this model the adaptive reading exceeds that control's by +1.19 logits (95% CI +0.85 to +1.54), positive in all ten seed pairs. Every item of the clean test lies inside the contaminated set, and an exposure set rebuilt only from simulated administrations reproduces the shift. Two further models show the same sign pattern; the tested selection-side defences lower the bias only at a clean-RMSE cost. Adaptive-test logs therefore need a threat model for their observers, and adaptive scores a whole-bank reading beside them.

Zenodo (CERN European Organization for Nuclear Research)
National Tsing Hua University (TW), North Carolina Exploring Cultural Heritage Online (US)
Quality Education
Psychometric Methodologies and Testing
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.