Do Complex Forecasting Models Beat Simple Benchmarks? A Small-Sample Case Study of Wheat, Corn, and Soybean Prices

Many studies report forecasting gains for sophisticated machine-learning and deep-learning models without testing whether those gains survive comparison against simple, reproducible benchmarks. This study tests that question for one-step-ahead monthly forecasting of wheat, corn, and soybean prices (FRED, 2000–2024; 300 observations, 179 training sequences). A 63-variable feature pipeline is built from MACD, RSI, Bollinger Bands, moving averages, volatility, momentum, and cross-commodity ratios. Eight learned models (SVR, Random Forest, Gradient Boosting, LSTM, GRU, Bidirectional LSTM, Attention-LSTM, and a MACD-Enhanced architecture, MADLF) are evaluated against two primary benchmarks, persistence and ARIMA (1, 1, 0), under one leakage-controlled, chronological protocol; four additional naive and statistical baselines (seasonal-naive, drift, ETS, Theta) are used for robustness. Forecast accuracy is assessed with RMSE, MAE, MASE, and Theil’s U2, and statistical significance with the Diebold–Mariano and Hansen’s SPA tests. No learned model shows statistically significant superiority over the naive benchmarks. For wheat and soybean, Diebold–Mariano tests reject all 16 learned-model comparisons; for corn, 4 of 8 are rejected against persistence and 5 of 8 against ARIMA, while Random Forest, BiLSTM, and MADLF are statistically indistinguishable from both, though with only 53 paired test observations this does not establish equivalence. Hansen’s SPA test cannot reject persistence as at least as good as the best learned model for any commodity. Descriptively, persistence attains R2=0.747, 0.738, and 0.662 for wheat, corn, and soybean, respectively, versus 0.313, 0.773, and 0.399 for the best learned model (computed from the same forecast series used for the significance tests throughout this article). Relative-skill metrics, a budget-constrained tuning check, three alternative training-window splits, a rolling-origin evaluation, and a matched-series multivariate (VAR) baseline all point to the same pattern, though the three splits and the rolling-origin check draw on a shared, overlapping 264-month timeline rather than independent samples, so they should be read as consistent rather than as independent confirmations. The recurrent architectures used here have 174,000 to 312,000 parameters but are trained on only 179 sequences; given this, the most reasonable interpretation of the null result is that simple benchmarks are hard to beat in a small-sample, monthly-data setting, not that architectural complexity is unhelpful in general. The findings are conditional on the specific feature representation, sample size, forecasting horizon, and experimental setting used in this study, and should not be read as a general claim that architectural complexity is unhelpful.

Authors

Institutions

Publication Details

Journal
Forecasting
Published
2026-09-25
DOI
https://doi.org/10.3390/forecast8050092
Primary Topic
Forecasting Techniques and Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Do Complex Forecasting Models Beat Simple Benchmarks? A Small-Sample Case Study of Wheat, Corn, and Soybean Prices

Dora Maria Sangermán-Jarquín, Belén Hernández Hernández, Sergio Ernesto Medina–Cuéllar, Juan Manuel Vargas-Canales et al.
Forecasting
Forecasting Techniques and Applications
article

Do Complex Forecasting Models Beat Simple Benchmarks? A Small-Sample Case Study of Wheat, Corn, and Soybean Prices

Dora Maria Sangermán-Jarquín, Belén Hernández Hernández, Sergio Ernesto Medina–Cuéllar, Juan Manuel Vargas-Canales, Juan Antonio Bautista, Benito Rodríguez Haros, Sergio Orozco Cirilo, Alberto Valdes Cobos
article en

Abstract

Many studies report forecasting gains for sophisticated machine-learning and deep-learning models without testing whether those gains survive comparison against simple, reproducible benchmarks. This study tests that question for one-step-ahead monthly forecasting of wheat, corn, and soybean prices (FRED, 2000–2024; 300 observations, 179 training sequences). A 63-variable feature pipeline is built from MACD, RSI, Bollinger Bands, moving averages, volatility, momentum, and cross-commodity ratios. Eight learned models (SVR, Random Forest, Gradient Boosting, LSTM, GRU, Bidirectional LSTM, Attention-LSTM, and a MACD-Enhanced architecture, MADLF) are evaluated against two primary benchmarks, persistence and ARIMA (1, 1, 0), under one leakage-controlled, chronological protocol; four additional naive and statistical baselines (seasonal-naive, drift, ETS, Theta) are used for robustness. Forecast accuracy is assessed with RMSE, MAE, MASE, and Theil’s U2, and statistical significance with the Diebold–Mariano and Hansen’s SPA tests. No learned model shows statistically significant superiority over the naive benchmarks. For wheat and soybean, Diebold–Mariano tests reject all 16 learned-model comparisons; for corn, 4 of 8 are rejected against persistence and 5 of 8 against ARIMA, while Random Forest, BiLSTM, and MADLF are statistically indistinguishable from both, though with only 53 paired test observations this does not establish equivalence. Hansen’s SPA test cannot reject persistence as at least as good as the best learned model for any commodity. Descriptively, persistence attains R2=0.747, 0.738, and 0.662 for wheat, corn, and soybean, respectively, versus 0.313, 0.773, and 0.399 for the best learned model (computed from the same forecast series used for the significance tests throughout this article). Relative-skill metrics, a budget-constrained tuning check, three alternative training-window splits, a rolling-origin evaluation, and a matched-series multivariate (VAR) baseline all point to the same pattern, though the three splits and the rolling-origin check draw on a shared, overlapping 264-month timeline rather than independent samples, so they should be read as consistent rather than as independent confirmations. The recurrent architectures used here have 174,000 to 312,000 parameters but are trained on only 179 sequences; given this, the most reasonable interpretation of the null result is that simple benchmarks are hard to beat in a small-sample, monthly-data setting, not that architectural complexity is unhelpful in general. The findings are conditional on the specific feature representation, sample size, forecasting horizon, and experimental setting used in this study, and should not be read as a general claim that architectural complexity is unhelpful.

ForecastingVol. 8(5)
Universidad de Guanajuato (MX), Universidad de Salamanca (ES), Instituto Nacional de Investigaciones Forestales Agrícolas y Pecuarias (MX), University of Celaya (MX)
Openalex Percentile: Top 7%
Forecasting Techniques and Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.