Benchmark study of ensemble machine learning with feature selection for rice yield prediction

Abstract Rice is a major staple food crop, and accurate yield prediction is essential for ensuring food security and supporting agricultural decision-making, especially under changing climatic conditions. In order to predict rice output utilizing meteorological, soil, and agronomic factors, this study suggests an ensemble machine learning approach. Standard performance metrics, including RMSE, MAE, and R 2 , were used to develop and assess several machine learning models, including Random Forest, XGBoost, LightGBM, and Gradient Boosting. The findings demonstrate that every single model performs well, obtaining R 2 values of more than 0.975, demonstrating their capacity to capture intricate nonlinear interactions in the dataset. Random Forest, XGBoost, Gradient Boosting, and LightGBM were combined to create an ensemble model based on a voting regressor in order to further enhance performance. To find the most pertinent predictors and improve model performance, sophisticated feature selection methods including Boruta and permutation significance were also used. The experimental findings show that the ensemble model with permutation-based feature selection performed best, with the highest accuracy (R 2 = 0.9787) and the lowest prediction error (RMSE = 5.4421, MAE = 4.3243). Five-fold cross-validation further confirmed the robustness and generalization capability of the proposed framework, achieving a mean CVRMSE of 5.496 ± 0.1944, a mean CVMAE of 4.3634 ± 0.1684, and a mean CV R 2 of 0.9789 ± 0.0018. The most significant elements influencing rice yield, according to feature importance analysis, are Fertilizer_Used_kg, Rainfall_mm, Temperature_C, and Humidity_pct. The results verify that prediction accuracy, resilience, and interpretability are slightly improved when ensemble learning and feature selection are combined. The proposed framework provides a data-driven approach for rice yield prediction using the evaluated synthetic dataset. The results demonstrate the effectiveness of combining ensemble learning and feature selection for improving prediction accuracy and model interpretability. It is important to highlight that this study assesses the proposed ensemble learning framework on a controlled synthetic dataset and serves as a benchmark for machine learning performance in rice yield prediction. Although the results show that ensemble models and feature selection improve prediction accuracy and interpretability, further validation with different real-world field datasets is required before applying the framework to operational agricultural decision-making.

Authors

Publication Details

Journal
Discover Artificial Intelligence
Published
2026-09-22
DOI
https://doi.org/10.1007/s44163-026-02201-3
Primary Topic
Smart Agriculture and AI
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Benchmark study of ensemble machine learning with feature selection for rice yield prediction

Fikadu Berie Adugna, Yonatan Motbaynor Tebabal, Esubalew Asmara Desta
Discover Artificial Intelligence
Smart Agriculture and AI
article

Benchmark study of ensemble machine learning with feature selection for rice yield prediction

Fikadu Berie Adugna, Yonatan Motbaynor Tebabal, Esubalew Asmara Desta
article en

Abstract

Abstract Rice is a major staple food crop, and accurate yield prediction is essential for ensuring food security and supporting agricultural decision-making, especially under changing climatic conditions. In order to predict rice output utilizing meteorological, soil, and agronomic factors, this study suggests an ensemble machine learning approach. Standard performance metrics, including RMSE, MAE, and R 2 , were used to develop and assess several machine learning models, including Random Forest, XGBoost, LightGBM, and Gradient Boosting. The findings demonstrate that every single model performs well, obtaining R 2 values of more than 0.975, demonstrating their capacity to capture intricate nonlinear interactions in the dataset. Random Forest, XGBoost, Gradient Boosting, and LightGBM were combined to create an ensemble model based on a voting regressor in order to further enhance performance. To find the most pertinent predictors and improve model performance, sophisticated feature selection methods including Boruta and permutation significance were also used. The experimental findings show that the ensemble model with permutation-based feature selection performed best, with the highest accuracy (R 2 = 0.9787) and the lowest prediction error (RMSE = 5.4421, MAE = 4.3243). Five-fold cross-validation further confirmed the robustness and generalization capability of the proposed framework, achieving a mean CVRMSE of 5.496 ± 0.1944, a mean CVMAE of 4.3634 ± 0.1684, and a mean CV R 2 of 0.9789 ± 0.0018. The most significant elements influencing rice yield, according to feature importance analysis, are Fertilizer_Used_kg, Rainfall_mm, Temperature_C, and Humidity_pct. The results verify that prediction accuracy, resilience, and interpretability are slightly improved when ensemble learning and feature selection are combined. The proposed framework provides a data-driven approach for rice yield prediction using the evaluated synthetic dataset. The results demonstrate the effectiveness of combining ensemble learning and feature selection for improving prediction accuracy and model interpretability. It is important to highlight that this study assesses the proposed ensemble learning framework on a controlled synthetic dataset and serves as a benchmark for machine learning performance in rice yield prediction. Although the results show that ensemble models and feature selection improve prediction accuracy and interpretability, further validation with different real-world field datasets is required before applying the framework to operational agricultural decision-making.

Discover Artificial IntelligenceVol. 6(1)
Zero hunger
Openalex Percentile: Top 13%
Smart Agriculture and AI
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.