Crop Yield Estimation with MODIS Derived Normalized Difference Vegetation Index and Comparative Study on Crop Yield Prediction Among Linear Regression, Random Forest and Gradient Boosting as Well as CatBoost

This paper presents the design, development, and evaluation of a machine-learning system built to forecast agricultural crop yields across Indian states between 2000 and 2026, together with a complementary, national-scale verification of predicted crop yield using a MODIS-derived NDVI time series (MOD13A3.061 Vegetation Indices Monthly L3 Global 1 km SIN Grid). Although many prior studies address crop-yield prediction with linear regression, random forest, gradient boosting, and related methods, a complementary, aggregate-level verification method for predicted crop yield has rarely been proposed. This article contributes such a method, together with a complementary NDVI-based estimation approach for total foodgrain output. Crop yield and MODIS-derived NDVI are strongly correlated (r = 0.84 for annual maximum NDVI; r = 0.78 for annual mean NDVI), and a simple regression of total foodgrains on annual maximum NDVI alone reaches R2 = 0.70. Four modeling approaches—linear regression, random forest, gradient boosting, and CatBoost—were built and compared using a chronology-preserving, expanding-window walk-forward validation procedure with a final, untouched 2024–2026 holdout, rather than a random split; a companion leakage check confirmed that reported production is almost algebraically identical to reported yield and therefore had to be excluded from the feature set. Random forest produced the most reliable and consistent forecasts, reaching a mean absolute percentage error (MAPE) of 11.4% and R2 = 0.982 on the final holdout, ahead of CatBoost (MAPE = 11.6%, R2 = 0.969) and gradient boosting (MAPE = 13.0%, R2 = 0.908), and substantially ahead of linear regression, which failed to generalize to the holdout period (R2 = −10.67); across the walk-forward folds preceding this holdout, however, the three tree ensembles were statistically indistinguishable. A four-configuration ablation study confirms that most of this performance gain is attributable to the inclusion of MODIS-derived NDVI rather than to model choice alone. Prediction error varies considerably by crop, from under 10% MAPE for major staples (rice, wheat, maize, sugarcane, moong) to well over 80% MAPE for several lower-volume crops (soyabean, garlic, Sunn hemp, tobacco, potato). The paper also documents two consequential data-quality findings—a near-perfect algebraic relationship between production and yield, and a structural administrative reporting gap in 2020—and closes with directions for future work, including higher-resolution satellite inputs, temporal deep-learning architectures, additional environmental covariates, and explainable-AI analysis of feature contributions.

Authors

Institutions

Publication Details

Journal
Remote Sensing
Published
2026-09-10
DOI
https://doi.org/10.3390/rs18183107
Primary Topic
Remote Sensing in Agriculture
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Crop Yield Estimation with MODIS Derived Normalized Difference Vegetation Index and Comparative Study on Crop Yield Prediction Among Linear Regression, Random Forest and Gradient Boosting as Well as CatBoost

Kohei Arai, Sara Sanwal
Remote Sensing
Remote Sensing in Agriculture
article

Crop Yield Estimation with MODIS Derived Normalized Difference Vegetation Index and Comparative Study on Crop Yield Prediction Among Linear Regression, Random Forest and Gradient Boosting as Well as CatBoost

Kohei Arai, Sara Sanwal
article en

Abstract

This paper presents the design, development, and evaluation of a machine-learning system built to forecast agricultural crop yields across Indian states between 2000 and 2026, together with a complementary, national-scale verification of predicted crop yield using a MODIS-derived NDVI time series (MOD13A3.061 Vegetation Indices Monthly L3 Global 1 km SIN Grid). Although many prior studies address crop-yield prediction with linear regression, random forest, gradient boosting, and related methods, a complementary, aggregate-level verification method for predicted crop yield has rarely been proposed. This article contributes such a method, together with a complementary NDVI-based estimation approach for total foodgrain output. Crop yield and MODIS-derived NDVI are strongly correlated (r = 0.84 for annual maximum NDVI; r = 0.78 for annual mean NDVI), and a simple regression of total foodgrains on annual maximum NDVI alone reaches R2 = 0.70. Four modeling approaches—linear regression, random forest, gradient boosting, and CatBoost—were built and compared using a chronology-preserving, expanding-window walk-forward validation procedure with a final, untouched 2024–2026 holdout, rather than a random split; a companion leakage check confirmed that reported production is almost algebraically identical to reported yield and therefore had to be excluded from the feature set. Random forest produced the most reliable and consistent forecasts, reaching a mean absolute percentage error (MAPE) of 11.4% and R2 = 0.982 on the final holdout, ahead of CatBoost (MAPE = 11.6%, R2 = 0.969) and gradient boosting (MAPE = 13.0%, R2 = 0.908), and substantially ahead of linear regression, which failed to generalize to the holdout period (R2 = −10.67); across the walk-forward folds preceding this holdout, however, the three tree ensembles were statistically indistinguishable. A four-configuration ablation study confirms that most of this performance gain is attributable to the inclusion of MODIS-derived NDVI rather than to model choice alone. Prediction error varies considerably by crop, from under 10% MAPE for major staples (rice, wheat, maize, sugarcane, moong) to well over 80% MAPE for several lower-volume crops (soyabean, garlic, Sunn hemp, tobacco, potato). The paper also documents two consequential data-quality findings—a near-perfect algebraic relationship between production and yield, and a structural administrative reporting gap in 2020—and closes with directions for future work, including higher-resolution satellite inputs, temporal deep-learning architectures, additional environmental covariates, and explainable-AI analysis of feature contributions.

Remote SensingVol. 18(18)
Saga University (JP), University of Prishtina (XK)
Zero hunger
Openalex Percentile: Top 10%
Remote Sensing in Agriculture
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.