A Library Update Moved the Threshold: Tracing One Changed Accept/Reject Decision to Float32 Scaling in scikit-learn

Updating a numerical library can change an individual accept/reject decision while the headline metrics, at the precision usually reported, stay the same. We document one such case and trace it to a single operation. We replayed an archived open-set microphone identification analysis from its stored float32 features. With scikit-learn 1.7.2, the version recorded for the original run, the replay reproduced all 202 compared fields of the original report bit for bit. With scikit-learn 1.8.0, 171 fields were identical, two differed in the last bits and 29 by more than 10-8; every pass/fail criterion stayed the same, and exactly one recording changed its decision. scikit-learn 1.9.1 gave the same 202 values as 1.8.0. The cause is a change in StandardScaler.transform for float32 input: version 1.8.0 casts the float64 mean and scale to float32 before subtracting and dividing; version 1.7.2 uses them in float64 and rounds each result once. Replacing only this arithmetic, in either direction, reproduced the other environment's score arrays exactly in all 24 folds. The decision changed because the threshold, a validation-score quantile, rose by 31 float32 ULP while the recording's own score fell by 7 ULP. On a second computer, in the three folds traced there, the same library change moved the same recording in the opposite direction. The change in scikit-learn is intentional and harmless for float64 data; the lesson concerns replication practice: when re-running a thresholded pipeline, compare individual decisions and the path that produces the threshold, not only aggregate metrics. Files: the preprint (PDF, with the implementation specification as Appendix A) and a reproduction package (ZIP): original analysis code and configuration (checked by SHA-256), stored float32 features of the 24 development phones, the historical report, pinned environments for scikit-learn 1.7.2, 1.8.0 and 1.9.1, replay and substitution scripts, our outputs, a NumPy-only reproducer of the mechanism, and VERIFIED_RESULTS.json with the script that rebuilds it from redacted copies of the saved outputs; README, LICENSE.md and a SHA-256 manifest. The JRC audio is not redistributed (doi:10.2905/JRC.FRE9H60). Licences: CC BY 4.0 for the text of the article and the result files; MIT for the code in reproduce/ and tools/. Stored features derived from the JRC dataset are reused under Commission Decision 2011/833/EU (see LICENSE.md in the ZIP). Funding: none. Competing interests: none. AI assistance: Claude (Anthropic) assisted with code, analysis scripts and drafting; Codex (OpenAI) and Gemini via Antigravity (Google) were used for independent review of the code, the numbers and the manuscript. The author is responsible for the content.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-08
DOI
https://doi.org/10.5281/zenodo.23224990
Primary Topic
Numerical Methods and Algorithms
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

A Library Update Moved the Threshold: Tracing One Changed Accept/Reject Decision to Float32 Scaling in scikit-learn

Mariusz Kulma
Zenodo (CERN European Organization for Nuclear Research)
Numerical Methods and Algorithms
preprint

A Library Update Moved the Threshold: Tracing One Changed Accept/Reject Decision to Float32 Scaling in scikit-learn

Mariusz Kulma
preprint en

Abstract

Updating a numerical library can change an individual accept/reject decision while the headline metrics, at the precision usually reported, stay the same. We document one such case and trace it to a single operation. We replayed an archived open-set microphone identification analysis from its stored float32 features. With scikit-learn 1.7.2, the version recorded for the original run, the replay reproduced all 202 compared fields of the original report bit for bit. With scikit-learn 1.8.0, 171 fields were identical, two differed in the last bits and 29 by more than 10-8; every pass/fail criterion stayed the same, and exactly one recording changed its decision. scikit-learn 1.9.1 gave the same 202 values as 1.8.0. The cause is a change in StandardScaler.transform for float32 input: version 1.8.0 casts the float64 mean and scale to float32 before subtracting and dividing; version 1.7.2 uses them in float64 and rounds each result once. Replacing only this arithmetic, in either direction, reproduced the other environment's score arrays exactly in all 24 folds. The decision changed because the threshold, a validation-score quantile, rose by 31 float32 ULP while the recording's own score fell by 7 ULP. On a second computer, in the three folds traced there, the same library change moved the same recording in the opposite direction. The change in scikit-learn is intentional and harmless for float64 data; the lesson concerns replication practice: when re-running a thresholded pipeline, compare individual decisions and the path that produces the threshold, not only aggregate metrics. Files: the preprint (PDF, with the implementation specification as Appendix A) and a reproduction package (ZIP): original analysis code and configuration (checked by SHA-256), stored float32 features of the 24 development phones, the historical report, pinned environments for scikit-learn 1.7.2, 1.8.0 and 1.9.1, replay and substitution scripts, our outputs, a NumPy-only reproducer of the mechanism, and VERIFIED_RESULTS.json with the script that rebuilds it from redacted copies of the saved outputs; README, LICENSE.md and a SHA-256 manifest. The JRC audio is not redistributed (doi:10.2905/JRC.FRE9H60). Licences: CC BY 4.0 for the text of the article and the result files; MIT for the code in reproduce/ and tools/. Stored features derived from the JRC dataset are reused under Commission Decision 2011/833/EU (see LICENSE.md in the ZIP). Funding: none. Competing interests: none. AI assistance: Claude (Anthropic) assisted with code, analysis scripts and drafting; Codex (OpenAI) and Gemini via Antigravity (Google) were used for independent review of the code, the numbers and the manuscript. The author is responsible for the content.

Zenodo (CERN European Organization for Nuclear Research)
Numerical Methods and Algorithms
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.