A Lightweight GDMM-YOLO11 Model for Cotton–Weed Instance Segmentation and Image-Plane Operation-Point Localization
To address the challenges of crop–weed instance segmentation and image-plane operation-point localization under visual similarity, background interference, leaf occlusion, and irregular plant morphology in cotton fields, a lightweight instance-segmentation model, GDMM-YOLO11, was developed. Based on YOLO11n-seg, the four backbone stage-transition downsampling convolutions at P2/4, P3/8, P4/16, and P5/32 were replaced with GhostConv to reduce redundant computation, a C3k2_DySnakeConv_Mona module was introduced to strengthen structural and multi-scale feature representation, and the original post-SPPF C2PSA block was replaced with mixed local channel attention (MLCA) to recalibrate high-level feature responses. Experiments were conducted on 2177 images containing 2856 annotated plant instances using a stratified 65%/15%/20% training–validation–test split. Across three independent runs with random seeds 3407, 3408, and 3409, GDMM-YOLO11 achieved a mask precision of 89.82 ± 2.60%, mask recall of 84.50 ± 2.49%, mask [email protected] of 89.76 ± 0.66%, and mask [email protected]:0.95 of 68.50 ± 0.34%. Relative to YOLO11n-seg, the corresponding three-run mean values were numerically higher by 3.52, 1.07, 1.67, and 2.33 percentage points, respectively, while the parameter count decreased from 2.836 M to 2.533 M and the computational cost decreased from 9.6 to 8.8 GFLOPs. For image-plane operation-point localization, 505 of 528 ground-truth weed instances obtained valid same-class mask matches. The proposed skeleton-constrained fused-center method achieved a ground-truth-mask inclusion rate of 99.41%, a conditional [email protected] of 91.29%, and an end-to-end [email protected] of 87.31%. Paired comparisons with the mask-centroid baseline showed statistically significant improvements in ground-truth-mask inclusion and weed-boundary clearance, whereas differences in localization error and [email protected]/0.15 were not statistically significant. TensorRT FP16 deployment on an NVIDIA Jetson AGX Orin achieved a model-only inference latency of 2.796 ± 0.115 ms, corresponding to 357.67 FPS. These results show that GDMM-YOLO11 provides a favorable accuracy–complexity trade-off while supporting image-plane operation-point generation and high-throughput model-only edge inference.
Authors
- Chenxu Zhao (ORCID: https://orcid.org/0000-0003-0714-8510)
- Yanhong Chen (ORCID: https://orcid.org/0000-0003-1795-0507)
- Yunjie Zhao (ORCID: https://orcid.org/0000-0001-5430-4390)
- Yongke Li
- Mengli Shan (ORCID: https://orcid.org/0009-0001-6561-0745)
- Lei Wang
Institutions
- Xinjiang Agricultural University (CN)
- Xinjiang University (CN)
Publication Details
- Journal
- Agronomy
- Published
- 2026-09-17
- DOI
- https://doi.org/10.3390/agronomy16181828
- Primary Topic
- Smart Agriculture and AI
- Type
- article
- Field-Weighted Citation Impact
- 0.00
Funders
- National Natural Science Foundation of China
- Science and Technology Department of Xinjiang Uyghur Autonomous Region