Ranking track records of unequal length: when the correction for selection bias helps, and when it costs
Marketplaces that rank strategies, vaults, funds or models almost always compare candidates whose track records differ in length, often by an order of magnitude. The standard corrections for selection bias do not admit this: the deflated Sharpe ratio deflates against the expected maximum of N trials of the same length, and the length weights used in production shrink a score after ranking rather than its variance before it. What the measurements say. Under a null in which nobody has skill, ranking by the Sharpe ratio hands first place to the shortest track record in 98.8 % of fields, and ranking by a deflated Sharpe score in 99.5 %. As a test the same deflation is conservative, 1.2 % for a nominal 5 %: its defect is not that it over-rejects but that it is used to rank. When one candidate genuinely is better and has a long record, the Sharpe ranking finds it 1.8 % of the time and the deflated score 0.8 %, against 73.8 % for the candidate's own t statistic and 87.2 % for a heteroscedastic empirical-Bayes posterior mean. And where it reverses. On 926 real products with histories of 30 to 1,677 days, the same length-aware rules lose. A short record there predicts the future as well as a long one, with rank correlations of +0.37 to +0.66 in every length bucket. The two regimes are separated by one measurable quantity, the ratio of the cross-sectional spread of true skill to the estimator's sampling variance: about 0.03 in the simulation and 561 to 18,936 in the real universe. The recommendation is therefore a diagnostic rather than an estimator. What can be published honestly: the length beside the score, the ratio above before a rule is chosen, and the set of candidates that cannot be separated from the winner, which in the real data is 20 to 52 of the 60 highest-ranked products. The work is pre-registered: the universe, the rules, the metrics and the ways it could fail were hash-stamped and committed before any outcome data was read. Two of the author's own assumptions did not survive the run and are reported as corrections, and one pre-registered metric failed on its own terms and is reported rather than replaced. The deposit contains the pre-registration, the estimators, the tests (including the two mistakes as regression tests), the simulations, the walk-forward, and the raw per-product histories.
Authors
- Ilpo Väätäinen
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-18
- DOI
- https://doi.org/10.5281/zenodo.22833947
- Primary Topic
- Meta-analysis and systematic reviews
- Type
- preprint