Freedman’s Paradox at Forty: A Practitioner’s Guide to Variable Selection Failures and Safeguards

In 1983, Freedman showed that variable selection on pure noise can produce spurious regressions. We benchmark 17 select-and-report workflows from five families across six Gaussian linear scenarios and add two reference-only procedures; four of the 19 provide valid-inference references. On common support, AIC stepwise falsely selects at least one variable in 89.9% of null samples. This risk does not fall as n grows to 5,000; larger samples improve signal recovery. No method family dominates the error–recovery curves across the evaluated cells and prespecified tuning paths. Unlike the search and screening defaults, the safeguards set an error target in advance. Apparent in-sample fit after selection often corresponds to negative out-of-sample R2. Among finite intervals, the valid-reference procedures’ native intervals achieve 94.8% to 95.1% marginal coverage. Under the complete null, conditional-on-selection coverage of naive intervals from the compared selection workflows is at most 7.8%. Workflow choice should therefore follow the goal: inference, prediction, or exploration.

Authors

Publication Details

Journal
The American Statistician
Published
2026-10-07
DOI
https://doi.org/10.1080/00031305.2026.2742880
Primary Topic
Statistical Methods and Inference
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Freedman’s Paradox at Forty: A Practitioner’s Guide to Variable Selection Failures and Safeguards

Yulia Vakulenko, Diogo Figueirinhas
The American Statistician
Statistical Methods and Inference
article

Freedman’s Paradox at Forty: A Practitioner’s Guide to Variable Selection Failures and Safeguards

Yulia Vakulenko, Diogo Figueirinhas
article en

Abstract

In 1983, Freedman showed that variable selection on pure noise can produce spurious regressions. We benchmark 17 select-and-report workflows from five families across six Gaussian linear scenarios and add two reference-only procedures; four of the 19 provide valid-inference references. On common support, AIC stepwise falsely selects at least one variable in 89.9% of null samples. This risk does not fall as n grows to 5,000; larger samples improve signal recovery. No method family dominates the error–recovery curves across the evaluated cells and prespecified tuning paths. Unlike the search and screening defaults, the safeguards set an error target in advance. Apparent in-sample fit after selection often corresponds to negative out-of-sample R2. Among finite intervals, the valid-reference procedures’ native intervals achieve 94.8% to 95.1% marginal coverage. Under the complete null, conditional-on-selection coverage of naive intervals from the compared selection workflows is at most 7.8%. Workflow choice should therefore follow the goal: inference, prediction, or exploration.

The American Statistician
Openalex Percentile: Top 11%
Statistical Methods and Inference
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Freedman’s Paradox at Forty: A Practitioner’s Guide to Variable Selection Failures and Safeguards — Yulia Vakulenko, Diogo Figueirinhas · The American Statistician (2026) | TGRS Research Map | TGRS