Do Complex Forecasting Models Beat Simple Benchmarks? A Small-Sample Case Study of Wheat, Corn, and Soybean Prices
Many studies report forecasting gains for sophisticated machine-learning and deep-learning models without testing whether those gains survive comparison against simple, reproducible benchmarks. This study tests that question for one-step-ahead monthly forecasting of wheat, corn, and soybean prices (FRED, 2000–2024; 300 observations, 179 training sequences). A 63-variable feature pipeline is built from MACD, RSI, Bollinger Bands, moving averages, volatility, momentum, and cross-commodity ratios. Eight learned models (SVR, Random Forest, Gradient Boosting, LSTM, GRU, Bidirectional LSTM, Attention-LSTM, and a MACD-Enhanced architecture, MADLF) are evaluated against two primary benchmarks, persistence and ARIMA (1, 1, 0), under one leakage-controlled, chronological protocol; four additional naive and statistical baselines (seasonal-naive, drift, ETS, Theta) are used for robustness. Forecast accuracy is assessed with RMSE, MAE, MASE, and Theil’s U2, and statistical significance with the Diebold–Mariano and Hansen’s SPA tests. No learned model shows statistically significant superiority over the naive benchmarks. For wheat and soybean, Diebold–Mariano tests reject all 16 learned-model comparisons; for corn, 4 of 8 are rejected against persistence and 5 of 8 against ARIMA, while Random Forest, BiLSTM, and MADLF are statistically indistinguishable from both, though with only 53 paired test observations this does not establish equivalence. Hansen’s SPA test cannot reject persistence as at least as good as the best learned model for any commodity. Descriptively, persistence attains R2=0.747, 0.738, and 0.662 for wheat, corn, and soybean, respectively, versus 0.313, 0.773, and 0.399 for the best learned model (computed from the same forecast series used for the significance tests throughout this article). Relative-skill metrics, a budget-constrained tuning check, three alternative training-window splits, a rolling-origin evaluation, and a matched-series multivariate (VAR) baseline all point to the same pattern, though the three splits and the rolling-origin check draw on a shared, overlapping 264-month timeline rather than independent samples, so they should be read as consistent rather than as independent confirmations. The recurrent architectures used here have 174,000 to 312,000 parameters but are trained on only 179 sequences; given this, the most reasonable interpretation of the null result is that simple benchmarks are hard to beat in a small-sample, monthly-data setting, not that architectural complexity is unhelpful in general. The findings are conditional on the specific feature representation, sample size, forecasting horizon, and experimental setting used in this study, and should not be read as a general claim that architectural complexity is unhelpful.
Authors
- Dora Maria Sangermán-Jarquín (ORCID: https://orcid.org/0000-0002-9658-1182)
- Belén Hernández Hernández
- Sergio Ernesto Medina–Cuéllar (ORCID: https://orcid.org/0000-0003-1883-935X)
- Juan Manuel Vargas-Canales (ORCID: https://orcid.org/0000-0003-1918-9395)
- Juan Antonio Bautista (ORCID: https://orcid.org/0000-0003-3326-7230)
- Benito Rodríguez Haros
- Sergio Orozco Cirilo
- Alberto Valdes Cobos
Institutions
- Universidad de Guanajuato (MX)
- Universidad de Salamanca (ES)
- Instituto Nacional de Investigaciones Forestales Agrícolas y Pecuarias (MX)
- University of Celaya (MX)
Publication Details
- Journal
- Forecasting
- Published
- 2026-09-25
- DOI
- https://doi.org/10.3390/forecast8050092
- Primary Topic
- Forecasting Techniques and Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00