A Library Update Moved the Threshold: Tracing One Changed Accept/Reject Decision to Float32 Scaling in scikit-learn
Updating a numerical library can change an individual accept/reject decision while the headline metrics, at the precision usually reported, stay the same. We document one such case and trace it to a single operation. We replayed an archived open-set microphone identification analysis from its stored float32 features. With scikit-learn 1.7.2, the version recorded for the original run, the replay reproduced all 202 compared fields of the original report bit for bit. With scikit-learn 1.8.0, 171 fields were identical, two differed in the last bits and 29 by more than 10-8; every pass/fail criterion stayed the same, and exactly one recording changed its decision. scikit-learn 1.9.1 gave the same 202 values as 1.8.0. The cause is a change in StandardScaler.transform for float32 input: version 1.8.0 casts the float64 mean and scale to float32 before subtracting and dividing; version 1.7.2 uses them in float64 and rounds each result once. Replacing only this arithmetic, in either direction, reproduced the other environment's score arrays exactly in all 24 folds. The decision changed because the threshold, a validation-score quantile, rose by 31 float32 ULP while the recording's own score fell by 7 ULP. On a second computer, in the three folds traced there, the same library change moved the same recording in the opposite direction. The change in scikit-learn is intentional and harmless for float64 data; the lesson concerns replication practice: when re-running a thresholded pipeline, compare individual decisions and the path that produces the threshold, not only aggregate metrics. Files: the preprint (PDF, with the implementation specification as Appendix A) and a reproduction package (ZIP): original analysis code and configuration (checked by SHA-256), stored float32 features of the 24 development phones, the historical report, pinned environments for scikit-learn 1.7.2, 1.8.0 and 1.9.1, replay and substitution scripts, our outputs, a NumPy-only reproducer of the mechanism, and VERIFIED_RESULTS.json with the script that rebuilds it from redacted copies of the saved outputs; README, LICENSE.md and a SHA-256 manifest. The JRC audio is not redistributed (doi:10.2905/JRC.FRE9H60). Licences: CC BY 4.0 for the text of the article and the result files; MIT for the code in reproduce/ and tools/. Stored features derived from the JRC dataset are reused under Commission Decision 2011/833/EU (see LICENSE.md in the ZIP). Funding: none. Competing interests: none. AI assistance: Claude (Anthropic) assisted with code, analysis scripts and drafting; Codex (OpenAI) and Gemini via Antigravity (Google) were used for independent review of the code, the numbers and the manuscript. The author is responsible for the content.
Authors
- Mariusz Kulma (ORCID: https://orcid.org/0009-0000-5550-8723)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-08
- DOI
- https://doi.org/10.5281/zenodo.23224990
- Primary Topic
- Numerical Methods and Algorithms
- Type
- preprint