The Null Zoo: Size and Power of Backtest-Overfitting Corrections on Synthetic Searches with Known Ground Truth

Corrections for backtest overfitting—the deflated Sharpe ratio, the multiple-testing haircut, family-wise adjustments and bootstrap tests—are derived under assumptions that real trading research violates. We introduce the Null Zoo, an open and versioned simulation benchmark that draws complete research searches (20 strategies of 504 daily returns) from eight return families with known ground truth and scores seven validators by size, raw power and size-adjusted power. A lag-one autocorrelation of 0.2 raises the false-positive rate of every test that ignores it, except the already near-zero deflated Sharpe ratio, to between 18.79% and 19.65% at a nominal 5%; deflating the t-statistic by Lo’s autocorrelation-adjusted annualisation factor brings it to 5.52% in this favourable AR(1) case, at a cost of a slight excess size elsewhere. Under negative skew the Student-t tests reject 7.60%; a non-normal standard error and an i.i.d. bootstrap each repair one shape and break another, and no validator is calibrated across skew and fat tails; under fat tails, selection favours series with a large positive outlier. Under i.i.d. normal returns the deflated Sharpe ratio with its authors’ 0.95 acceptance rule fails to accept 90.5% of searches containing a strategy with a true Sharpe ratio of 2, because the rule tests a harder hypothesis than no skill and the skilled strategy raises its own benchmark; size-adjusted, it recovers most of that power (46.9% against 52.6%) but trails the Šidák test by 5.2–18.0 points across families. A bootstrap’s number of resamples materially affects its power. The published results reproduce byte for byte from the released code, and a robustness run with a different generator, 10,000 replications per cell and a finer bootstrap shows the findings are not artefacts of the generator or of Monte Carlo error. Code, result files and LaTeX source: https://github.com/arhancanli/null-zoo (release v1.0.0, MIT licence). Every number in the paper is generated from the released result files, and the published benchmark reproduces byte for byte in continuous integration.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-28
DOI
https://doi.org/10.5281/zenodo.23018988
Primary Topic
Sports Analytics and Performance
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

The Null Zoo: Size and Power of Backtest-Overfitting Corrections on Synthetic Searches with Known Ground Truth

Arhan Canli
Zenodo (CERN European Organization for Nuclear Research)
Sports Analytics and Performance
preprint

The Null Zoo: Size and Power of Backtest-Overfitting Corrections on Synthetic Searches with Known Ground Truth

Arhan Canli
preprint en

Abstract

Corrections for backtest overfitting—the deflated Sharpe ratio, the multiple-testing haircut, family-wise adjustments and bootstrap tests—are derived under assumptions that real trading research violates. We introduce the Null Zoo, an open and versioned simulation benchmark that draws complete research searches (20 strategies of 504 daily returns) from eight return families with known ground truth and scores seven validators by size, raw power and size-adjusted power. A lag-one autocorrelation of 0.2 raises the false-positive rate of every test that ignores it, except the already near-zero deflated Sharpe ratio, to between 18.79% and 19.65% at a nominal 5%; deflating the t-statistic by Lo’s autocorrelation-adjusted annualisation factor brings it to 5.52% in this favourable AR(1) case, at a cost of a slight excess size elsewhere. Under negative skew the Student-t tests reject 7.60%; a non-normal standard error and an i.i.d. bootstrap each repair one shape and break another, and no validator is calibrated across skew and fat tails; under fat tails, selection favours series with a large positive outlier. Under i.i.d. normal returns the deflated Sharpe ratio with its authors’ 0.95 acceptance rule fails to accept 90.5% of searches containing a strategy with a true Sharpe ratio of 2, because the rule tests a harder hypothesis than no skill and the skilled strategy raises its own benchmark; size-adjusted, it recovers most of that power (46.9% against 52.6%) but trails the Šidák test by 5.2–18.0 points across families. A bootstrap’s number of resamples materially affects its power. The published results reproduce byte for byte from the released code, and a robustness run with a different generator, 10,000 replications per cell and a finer bootstrap shows the findings are not artefacts of the generator or of Monte Carlo error. Code, result files and LaTeX source: https://github.com/arhancanli/null-zoo (release v1.0.0, MIT licence). Every number in the paper is generated from the released result files, and the published benchmark reproduces byte for byte in continuous integration.

Zenodo (CERN European Organization for Nuclear Research)
Peace, Justice and strong institutions
Sports Analytics and Performance
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.