AgriViT-LLM: A Dual-Head Vision Transformer Coupled with a GRPO-Trained Retrieval-Augmented Language Model for Crop Disease Diagnosis and Integrated Pest Management
Diagnosing a crop disease is only the first step; farmers also require clear, trustworthy guidance on how to treat it. This paper presents AgriViT-LLM, a two-stage system that first classifies the disease and estimates its severity using a dual-head Vision Transformer (DualHeadViT), then generates Integrated Pest Management (IPM) treatment recommendations using a retrieval-augmented Large Language Model (LLM). The treatment advice is grounded in a curated set of documents from the Indian Council of Agricultural Research (ICAR) and Acharya N. G. Ranga Agricultural University (ANGRAU), so recommendations can be traced back to a real source rather than the model's own guesswork. The classifier is trained on 49,144 images covering 46 in-distribution disease classes drawn from three public datasets (PlantVillage, Cotton, Banana) plus a field-image set (PlantDoc); three additional PlantVillage classes (Citrus HLB, Grape Esca, Tomato YLCV) are deliberately withheld to test how the system handles diseases it has never seen (out-of-distribution, or OOD). DualHeadViT reaches 99.10% accuracy on the held-out test set. The treatment module fine-tunes a Qwen2.5-7B language model, first with supervised fine-tuning and then with Group Relative Policy Optimization (GRPO) against a transparent, six-part reward, reasoning over evidence retrieved with a hybrid BM25+FAISS search and crop-aware re-ranking (a Search-R1-style approach). Controlled experiments show that both retrieval and GRPO matter: removing retrieval causes treatment quality to collapse on diseases the model has not seen before, and GRPO training raises the IPM score from 38.3 to 62.3 compared with supervised fine-tuning alone, mainly by making the recommendations more complete. Ablating architectural components (the dual head, severity head, and MixUp augmentation) makes little difference to already-saturated in-distribution accuracy, but each one consistently helps the model generalise better to real field photos. Overall, the system produces verifiable, well-sourced bio-pesticide, chemical, and cultural treatment recommendations — closing a gap left by systems that only classify the disease.
Authors
- Dr.N.P.Lavanya Kumari Varadhi Aakanksha
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-06
- DOI
- https://doi.org/10.5281/zenodo.23192796
- Primary Topic
- Smart Agriculture and AI
- Type
- article
- Field-Weighted Citation Impact
- 0.00