Crop Yield Estimation with MODIS Derived Normalized Difference Vegetation Index and Comparative Study on Crop Yield Prediction Among Linear Regression, Random Forest and Gradient Boosting as Well as CatBoost
This paper presents the design, development, and evaluation of a machine-learning system built to forecast agricultural crop yields across Indian states between 2000 and 2026, together with a complementary, national-scale verification of predicted crop yield using a MODIS-derived NDVI time series (MOD13A3.061 Vegetation Indices Monthly L3 Global 1 km SIN Grid). Although many prior studies address crop-yield prediction with linear regression, random forest, gradient boosting, and related methods, a complementary, aggregate-level verification method for predicted crop yield has rarely been proposed. This article contributes such a method, together with a complementary NDVI-based estimation approach for total foodgrain output. Crop yield and MODIS-derived NDVI are strongly correlated (r = 0.84 for annual maximum NDVI; r = 0.78 for annual mean NDVI), and a simple regression of total foodgrains on annual maximum NDVI alone reaches R2 = 0.70. Four modeling approaches—linear regression, random forest, gradient boosting, and CatBoost—were built and compared using a chronology-preserving, expanding-window walk-forward validation procedure with a final, untouched 2024–2026 holdout, rather than a random split; a companion leakage check confirmed that reported production is almost algebraically identical to reported yield and therefore had to be excluded from the feature set. Random forest produced the most reliable and consistent forecasts, reaching a mean absolute percentage error (MAPE) of 11.4% and R2 = 0.982 on the final holdout, ahead of CatBoost (MAPE = 11.6%, R2 = 0.969) and gradient boosting (MAPE = 13.0%, R2 = 0.908), and substantially ahead of linear regression, which failed to generalize to the holdout period (R2 = −10.67); across the walk-forward folds preceding this holdout, however, the three tree ensembles were statistically indistinguishable. A four-configuration ablation study confirms that most of this performance gain is attributable to the inclusion of MODIS-derived NDVI rather than to model choice alone. Prediction error varies considerably by crop, from under 10% MAPE for major staples (rice, wheat, maize, sugarcane, moong) to well over 80% MAPE for several lower-volume crops (soyabean, garlic, Sunn hemp, tobacco, potato). The paper also documents two consequential data-quality findings—a near-perfect algebraic relationship between production and yield, and a structural administrative reporting gap in 2020—and closes with directions for future work, including higher-resolution satellite inputs, temporal deep-learning architectures, additional environmental covariates, and explainable-AI analysis of feature contributions.
Authors
- Kohei Arai (ORCID: https://orcid.org/0009-0001-6433-1592)
- Sara Sanwal (ORCID: https://orcid.org/0009-0003-7944-0948)
Institutions
- Saga University (JP)
- University of Prishtina (XK)
Publication Details
- Journal
- Remote Sensing
- Published
- 2026-09-10
- DOI
- https://doi.org/10.3390/rs18183107
- Primary Topic
- Remote Sensing in Agriculture
- Type
- article
- Field-Weighted Citation Impact
- 0.00