A global-local hybrid vision-language framework for interpretable grapevine disease diagnosis
Abstract Plant diseases pose a severe threat to agricultural productivity and global food security. Although Convolutional Neural Networks have achieved impressive classification accuracy in controlled environments, they often lack semantic reasoning capabilities and operate as uninterpretable “black boxes.” Recently, Vision-Language Models have emerged as powerful multimodal reasoning systems. However, they suffer from the “attention sink” phenomenon when applied to complex real-world images, wasting computational resources on background noise rather than microscopic necrotic lesions. In this paper, we propose a Global–Local Hybrid Framework that mimics the diagnostic workflow of agricultural experts. The system leverages an unsupervised Saliency Auto-Crop algorithm within the CIE-LAB color space to automatically extract local disease patches without the need for bounding box annotations. A dual image stream, comprising global context and local patches, is fed into LLaVA−1.5 7B and fine-tuned using 4-bit QLoRA to ensure parameter efficiency. Furthermore, we extract Cross-Attention Heatmaps from the final layer of the decoder to provide Explainable Artificial Intelligence capabilities. Experimental results on a grapevine disease dataset demonstrate that our method achieves an overall accuracy of 98.7%, significantly outperforming zero-shot baselines while providing explicit semantic interpretability for precision agriculture.
Authors
- Bùi Thanh Hùng (ORCID: https://orcid.org/0000-0002-9400-7582)
- Lang Hoang Son
Institutions
- Industrial University of Ho Chi Minh City (VN)
- Ho Chi Minh City University of Science (VN)
Publication Details
- Journal
- Discover Artificial Intelligence
- Published
- 2026-10-07
- DOI
- https://doi.org/10.1007/s44163-026-02463-x
- Primary Topic
- Smart Agriculture and AI
- Type
- article
- Field-Weighted Citation Impact
- 0.00