Lesion-Aware Multiple-Instance Learning for Grapevine Disease Recognition from Field RGB Images

This retrospective offline study examines how dense manual regional annotations can be organized for image-level grapevine disease recognition using pre-acquired HERMOS RGB images. After audit and filtering, 491 source images and 13,765 Pascal VOC regions were retained. A leakage-safe source-image split was used across full-image classification, regional crop classification, crop-to-image aggregation, annotation-guided multiple-instance learning (MIL), and YOLO11 localization. MIL bags in training, validation, and testing were constructed from crops derived from manual boxes; the results therefore describe annotation-assisted aggregation rather than raw-image deployment. The original ConvNeXt-Tiny MIL model achieved macro-F1 0.922 (95% source-image bootstrap CI 0.882–0.956), compared with 0.836 (0.754–0.895) for full-image ConvNeXt-Tiny and 0.869 (0.804–0.917) for the best fixed crop aggregation. Paired macro-F1 differences favored MIL over these baselines by 0.087 (95% CI 0.029–0.163; Holm-adjusted p = 0.013) and 0.053 (0.005–0.114; adjusted p = 0.036), respectively. Downy mildew precision and recall were 1.000 on nine positive test images, but the corresponding exact 95% binomial intervals were 0.664–1.000. Controlled ablations produced lower point estimates without learned attention or the auxiliary instance loss, whereas removing context expansion did not reduce performance; none of the three paired ablation tests was significant after Holm correction. Increasing detector capacity from YOLO11n to YOLO11s did not improve full-image [email protected] (0.175 versus 0.170). These findings support annotation-guided local-evidence aggregation as an offline exploratory benchmark, while not establishing split-to-split stability, computational efficiency, or external-vineyard generalization.

Authors

Institutions

Publication Details

Journal
Sensors
Published
2026-10-07
DOI
https://doi.org/10.3390/s26196318
Primary Topic
Smart Agriculture and AI
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Lesion-Aware Multiple-Instance Learning for Grapevine Disease Recognition from Field RGB Images

Tuğba Özacar, V. Ozacar, Övünç Öztürk, Bora Canbula
Sensors
Smart Agriculture and AI
article

Lesion-Aware Multiple-Instance Learning for Grapevine Disease Recognition from Field RGB Images

Tuğba Özacar, V. Ozacar, Övünç Öztürk, Bora Canbula
article en

Abstract

This retrospective offline study examines how dense manual regional annotations can be organized for image-level grapevine disease recognition using pre-acquired HERMOS RGB images. After audit and filtering, 491 source images and 13,765 Pascal VOC regions were retained. A leakage-safe source-image split was used across full-image classification, regional crop classification, crop-to-image aggregation, annotation-guided multiple-instance learning (MIL), and YOLO11 localization. MIL bags in training, validation, and testing were constructed from crops derived from manual boxes; the results therefore describe annotation-assisted aggregation rather than raw-image deployment. The original ConvNeXt-Tiny MIL model achieved macro-F1 0.922 (95% source-image bootstrap CI 0.882–0.956), compared with 0.836 (0.754–0.895) for full-image ConvNeXt-Tiny and 0.869 (0.804–0.917) for the best fixed crop aggregation. Paired macro-F1 differences favored MIL over these baselines by 0.087 (95% CI 0.029–0.163; Holm-adjusted p = 0.013) and 0.053 (0.005–0.114; adjusted p = 0.036), respectively. Downy mildew precision and recall were 1.000 on nine positive test images, but the corresponding exact 95% binomial intervals were 0.664–1.000. Controlled ablations produced lower point estimates without learned attention or the auxiliary instance loss, whereas removing context expansion did not reduce performance; none of the three paired ablation tests was significant after Holm correction. Increasing detector capacity from YOLO11n to YOLO11s did not improve full-image [email protected] (0.175 versus 0.170). These findings support annotation-guided local-evidence aggregation as an offline exploratory benchmark, while not establishing split-to-split stability, computational efficiency, or external-vineyard generalization.

SensorsVol. 26(19)
Manisa Celal Bayar University (TR), Dokuz Eylül University (TR)
Openalex Percentile: Top 14%
Smart Agriculture and AI
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.