Development of Artificial Intelligence Based Model to Predict Crop Yields of Biofertilizers
The research paper proposes the development and the analysis of an artificial intelligence model that can be used for predicting the crop yield under various biofertilizer treatment, types of soils and environmental conditions. The problem statement is the necessity of the prediction of the efficiency of biofertilizers taking into account the interaction of microbial strains, soil properties, inoculation rate, nutrients, crops, and climatic zones. In this research, there is a use of the dataset of 5,500 observations, including 500 real observations obtained during the published materials with biofertilizers and 5,000 artificial observations created by conditional tabular generative adversarial network (CTGAN) in python. The dataset contains data about maize and mixed crops, including such features as soil pH, nitrogen, phosphorus, potassium, soil moisture, biofertilizer type, inoculation dose, colony-forming units of microbes, soil type, climatic zone, and crop yield. There was used data pre-processing, including cleaning, encoding, scaling, and splitting the dataset into the training, validation, and test sets. In this research, such algorithms as linear regression, random forest, gradient boosting, and XG-boost were applied. The choice of models was based on the metrics R², RMSE, MAE, MAPE, and Nash-Sutcliffe Efficiency (NSE). The optimal algorithm for the maize dataset was XGBoost with the following parameters: R² = 0.9479, RMSE = 300.06 kg/ha, and MAE = 225.85 kg/ha. The most accurate prediction of the crop yield for the dataset of mixed crops was the result of the Random Forest algorithm: R² = 0.9990, RMSE = 234.91 kg/ha, and MAE = 113.73 kg/ha. The findings suggest that the interaction of Azotobacter and Azospirillum led to an increase of 28.7% in maize production in comparison with the control group. Variables that have the greatest influence on the process were found to be biofertilizer application, soil nitrogen, inoculation level, soil, and climate zone. Further, the best-performing model underwent optimization through feature elimination, quantization, and reduced precision, making it embeddable on the device. The optimization resulted in R² value equal to 0.9108, or about 96% of the original model’s goodness-of-fit measure. It can be seen that tree-based machine learning models can learn intricate relationships between biofertilizer application and yield. However, more testing with independent data is needed before the use of such models for practice.
Authors
- Murtala Aminu Baba
- Surajudeen Abdulsalam
- Anas Musa Shehu
Institutions
- Abubakar Tafawa Balewa University (NG)
Publication Details
- Journal
- Iconic Research and Engineering Journals
- Published
- 2026-09-29
- DOI
- https://doi.org/10.64388/irev10i3-1723480
- Primary Topic
- Smart Agriculture and AI
- Type
- article
- Field-Weighted Citation Impact
- 0.00