Human Medians and a Language-Model Forecast: Encompassing and Combination on One ForecastBench Round

I use the one ForecastBench round with individual-level human forecasts (July 2024) to compare two human medians, from 40 superforecaster ids and 500 public forecasters, with the best single language model, chosen on questions no human forecast. The humans did not see the model and could look things up; the model had no news. In preregistered tests the model encompassed neither median (coefficients 1.79 and 0.97), the average superforecaster forecast beat it by 3.4 Brier points ×100, and the average public forecast lost to it by 7.8. In analyses specified after these results, with readings locked before computation, fitted combinations gave the superforecaster median weight 1 and equal-weight pools were worse than that median alone, while the gain from combining the public median with the model on held-out questions had an interval that included zero. The data come from one round, the superforecasters had a group stage, and no news-augmented model variant beat the benchmark, so these results are not evidence about judgment. Version note. This version replaces "Where the Benchmark Can Err: Human Deviations from Model Forecasts and Their Return" (v3). The study, data, and preregistered tests are the same. The framing changed from human deviation to forecast encompassing and combination; the chess comparison and hypotheses H3 to H5 moved to a companion methods paper; and three post-results analyses (encompassing by source, out-of-sample combination, news-augmented variants) were added under SPEC v3, with readings locked before computation.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-05
DOI
https://doi.org/10.5281/zenodo.23147918
Primary Topic
Forecasting Techniques and Applications
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Human Medians and a Language-Model Forecast: Encompassing and Combination on One ForecastBench Round

Aaron Sun
Zenodo (CERN European Organization for Nuclear Research)
Forecasting Techniques and Applications
preprint

Human Medians and a Language-Model Forecast: Encompassing and Combination on One ForecastBench Round

Aaron Sun
preprint en

Abstract

I use the one ForecastBench round with individual-level human forecasts (July 2024) to compare two human medians, from 40 superforecaster ids and 500 public forecasters, with the best single language model, chosen on questions no human forecast. The humans did not see the model and could look things up; the model had no news. In preregistered tests the model encompassed neither median (coefficients 1.79 and 0.97), the average superforecaster forecast beat it by 3.4 Brier points ×100, and the average public forecast lost to it by 7.8. In analyses specified after these results, with readings locked before computation, fitted combinations gave the superforecaster median weight 1 and equal-weight pools were worse than that median alone, while the gain from combining the public median with the model on held-out questions had an interval that included zero. The data come from one round, the superforecasters had a group stage, and no news-augmented model variant beat the benchmark, so these results are not evidence about judgment. Version note. This version replaces "Where the Benchmark Can Err: Human Deviations from Model Forecasts and Their Return" (v3). The study, data, and preregistered tests are the same. The framing changed from human deviation to forecast encompassing and combination; the chess comparison and hypotheses H3 to H5 moved to a companion methods paper; and three post-results analyses (encompassing by source, out-of-sample combination, news-augmented variants) were added under SPEC v3, with readings locked before computation.

Zenodo (CERN European Organization for Nuclear Research)
Forecasting Techniques and Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Human Medians and a Language-Model Forecast: Encompassing and Combination on One ForecastBench Round — Aaron Sun · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS