Benchmark study of ensemble machine learning with feature selection for rice yield prediction
Abstract Rice is a major staple food crop, and accurate yield prediction is essential for ensuring food security and supporting agricultural decision-making, especially under changing climatic conditions. In order to predict rice output utilizing meteorological, soil, and agronomic factors, this study suggests an ensemble machine learning approach. Standard performance metrics, including RMSE, MAE, and R 2 , were used to develop and assess several machine learning models, including Random Forest, XGBoost, LightGBM, and Gradient Boosting. The findings demonstrate that every single model performs well, obtaining R 2 values of more than 0.975, demonstrating their capacity to capture intricate nonlinear interactions in the dataset. Random Forest, XGBoost, Gradient Boosting, and LightGBM were combined to create an ensemble model based on a voting regressor in order to further enhance performance. To find the most pertinent predictors and improve model performance, sophisticated feature selection methods including Boruta and permutation significance were also used. The experimental findings show that the ensemble model with permutation-based feature selection performed best, with the highest accuracy (R 2 = 0.9787) and the lowest prediction error (RMSE = 5.4421, MAE = 4.3243). Five-fold cross-validation further confirmed the robustness and generalization capability of the proposed framework, achieving a mean CVRMSE of 5.496 ± 0.1944, a mean CVMAE of 4.3634 ± 0.1684, and a mean CV R 2 of 0.9789 ± 0.0018. The most significant elements influencing rice yield, according to feature importance analysis, are Fertilizer_Used_kg, Rainfall_mm, Temperature_C, and Humidity_pct. The results verify that prediction accuracy, resilience, and interpretability are slightly improved when ensemble learning and feature selection are combined. The proposed framework provides a data-driven approach for rice yield prediction using the evaluated synthetic dataset. The results demonstrate the effectiveness of combining ensemble learning and feature selection for improving prediction accuracy and model interpretability. It is important to highlight that this study assesses the proposed ensemble learning framework on a controlled synthetic dataset and serves as a benchmark for machine learning performance in rice yield prediction. Although the results show that ensemble models and feature selection improve prediction accuracy and interpretability, further validation with different real-world field datasets is required before applying the framework to operational agricultural decision-making.
Authors
- Fikadu Berie Adugna (ORCID: https://orcid.org/0000-0001-7507-3982)
- Yonatan Motbaynor Tebabal
- Esubalew Asmara Desta
Publication Details
- Journal
- Discover Artificial Intelligence
- Published
- 2026-09-22
- DOI
- https://doi.org/10.1007/s44163-026-02201-3
- Primary Topic
- Smart Agriculture and AI
- Type
- article
- Field-Weighted Citation Impact
- 0.00