One Percent of the Bank: Adaptive Item Selection Reads Targeted Contamination Up While the Whole Bank Barely Moves

Adaptive testing is entering language-model evaluation as an efficiency device, and its low item exposure is offered as incidental contamination protection (Li et al., 2026b). That argument assumes untargeted leakage, yet a maximum-information test asks predictable items that anyone observing administrations can learn. Here we fine-tune a 1.5B-parameter model on 1% of a self-calibrated MMLU bank, chosen by exposure in a clean adaptive run, and score adaptive, random and whole-bank protocols on the same response vectors. A 41-item adaptive test then rises by 1.47 to 1.59 logits over its own clean reading (1.63 to 1.75 above the clean whole-bank reading) while the whole bank reads slightly lower. Because this concentrated test also moves under untargeted fine-tuning drift, we compare it with a random-item set trained to approximately the same loss of clean-item accuracy: on this model the adaptive reading exceeds that control's by +1.40 logits (95% CI +1.10 to +1.69), positive in all ten seed pairs. Every item of the clean test lies inside the contaminated set, and an exposure set rebuilt only from simulated administrations reproduces the shift. Two further models show the same sign pattern; none of the selection-side defences tested in simulation meets a joint target for bias and clean RMSE. In the released logs of two adaptive LLM evaluators, the most-administered 1% of each instrument's largest bank takes 30–40% (ATLAS, HellaSwag) and 60% (Fluid, MMLU, 100-item tests) of all administrations; across all their benchmarks (Fluid at 100 items) the most-administered 1% takes 5–60%. No change confined to a targeted set can move the expected reading of a reader that is design-unbiased for bank accuracy by more than the set's share of the bank: that expectation changes by exactly the change in bank accuracy. Adaptive-test logs therefore need a threat model for their observers, and adaptive scores a whole-bank reading beside them.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-05
DOI
https://doi.org/10.5281/zenodo.22709286
Primary Topic
Adversarial Robustness in Machine Learning
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

One Percent of the Bank: Adaptive Item Selection Reads Targeted Contamination Up While the Whole Bank Barely Moves

Ya-Fen Yeh, Guan-Yuan Chen
Zenodo (CERN European Organization for Nuclear Research)
Adversarial Robustness in Machine Learning
preprint

One Percent of the Bank: Adaptive Item Selection Reads Targeted Contamination Up While the Whole Bank Barely Moves

Ya-Fen Yeh, Guan-Yuan Chen
preprint en

Abstract

Adaptive testing is entering language-model evaluation as an efficiency device, and its low item exposure is offered as incidental contamination protection (Li et al., 2026b). That argument assumes untargeted leakage, yet a maximum-information test asks predictable items that anyone observing administrations can learn. Here we fine-tune a 1.5B-parameter model on 1% of a self-calibrated MMLU bank, chosen by exposure in a clean adaptive run, and score adaptive, random and whole-bank protocols on the same response vectors. A 41-item adaptive test then rises by 1.47 to 1.59 logits over its own clean reading (1.63 to 1.75 above the clean whole-bank reading) while the whole bank reads slightly lower. Because this concentrated test also moves under untargeted fine-tuning drift, we compare it with a random-item set trained to approximately the same loss of clean-item accuracy: on this model the adaptive reading exceeds that control's by +1.40 logits (95% CI +1.10 to +1.69), positive in all ten seed pairs. Every item of the clean test lies inside the contaminated set, and an exposure set rebuilt only from simulated administrations reproduces the shift. Two further models show the same sign pattern; none of the selection-side defences tested in simulation meets a joint target for bias and clean RMSE. In the released logs of two adaptive LLM evaluators, the most-administered 1% of each instrument's largest bank takes 30–40% (ATLAS, HellaSwag) and 60% (Fluid, MMLU, 100-item tests) of all administrations; across all their benchmarks (Fluid at 100 items) the most-administered 1% takes 5–60%. No change confined to a targeted set can move the expected reading of a reader that is design-unbiased for bank accuracy by more than the set's share of the bank: that expectation changes by exactly the change in bank accuracy. Adaptive-test logs therefore need a threat model for their observers, and adaptive scores a whole-bank reading beside them.

Zenodo (CERN European Organization for Nuclear Research)
National Tsing Hua University (TW), North Carolina Exploring Cultural Heritage Online (US)
Adversarial Robustness in Machine Learning
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.