Human Medians and a Language-Model Forecast: Encompassing and Combination on One ForecastBench Round
I use the one ForecastBench round with individual-level human forecasts (July 2024) to compare two human medians, from 40 superforecaster ids and 500 public forecasters, with the best single language model, chosen on questions no human forecast. The humans did not see the model and could look things up; the model had no news. In preregistered tests the model encompassed neither median (coefficients 1.79 and 0.97), the average superforecaster forecast beat it by 3.4 Brier points ×100, and the average public forecast lost to it by 7.8. In analyses specified after these results, with readings locked before computation, fitted combinations gave the superforecaster median weight 1 and equal-weight pools were worse than that median alone, while the gain from combining the public median with the model on held-out questions had an interval that included zero. The data come from one round, the superforecasters had a group stage, and no news-augmented model variant beat the benchmark, so these results are not evidence about judgment. Version note. This version replaces "Where the Benchmark Can Err: Human Deviations from Model Forecasts and Their Return" (v3). The study, data, and preregistered tests are the same. The framing changed from human deviation to forecast encompassing and combination; the chess comparison and hypotheses H3 to H5 moved to a companion methods paper; and three post-results analyses (encompassing by source, out-of-sample combination, news-augmented variants) were added under SPEC v3, with readings locked before computation.
Authors
- Aaron Sun (ORCID: https://orcid.org/0009-0000-2412-4727)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-05
- DOI
- https://doi.org/10.5281/zenodo.23147918
- Primary Topic
- Forecasting Techniques and Applications
- Type
- preprint