Apple Counting and Yield Estimation in High-Density Dwarf Orchards from Monocular UAV Videos Using Global-Motion-Compensated DeepSORT

Fruit occlusion, dense clustering, platform vibration, and cross-frame re-identification errors reduce the reliability of video-stream yield estimation in high-density dwarf apple orchards. To address these limitations, this study developed a vision-based framework for apple detection, tracking, dynamic counting, and yield estimation. Apple detection was improved by embedding a Shuffle Attention (SA) module into YOLOv12n and introducing SlideLoss for hard-sample reweighting, yielding the YOLOv12n-SA-SlideLoss detector. Following detection, the DeepSORT tracking module was also improved by augmenting it with global motion compensation (GMC) and a dual-criteria association strategy based on motion consistency and appearance similarity. Yield estimation was formulated as a counting-based compensation model using dynamic fruit counts, mean fruit mass, and an occlusion compensation factor, and was compared with a pixel-geometric regression model based on detection boxes. Field experiments were conducted in a high-density dwarf ‘Fuji’ apple orchard in Aksu, Xinjiang, China. YOLOv12n-SA-SlideLoss achieved an [email protected] of 95.7%, exceeding the YOLOv12n baseline by 2.3%. The improved DeepSORT reached 87.4% multiple object tracking accuracy and 86.7% multiple object tracking precision, while reducing identification switches by 22.2% relative to the original algorithm. The dynamic counting framework increased average counting accuracy from 85.6% to 93.5%. For yield estimation, the counting-based compensation model achieved a mean prediction accuracy of 85.2%, an RMSE of 18.5 kg, and a mean bias of −17.2 kg across six validation samples (n = 6), each covering multiple trees. Overall, the results indicate that combining detection, motion-compensated tracking, and counting-based correction improves stability of single-view video-stream yield estimation in high-density dwarf apple orchards.

Authors

Institutions

Publication Details

Journal
Agriculture
Published
2026-09-29
DOI
https://doi.org/10.3390/agriculture16192109
Primary Topic
Smart Agriculture and AI
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Apple Counting and Yield Estimation in High-Density Dwarf Orchards from Monocular UAV Videos Using Global-Motion-Compensated DeepSORT

Dae-Hyun Lee, Longsheng Fu, Bryan Gilbert Murengami, Jun Chen et al.
Agriculture
Smart Agriculture and AI
article

Apple Counting and Yield Estimation in High-Density Dwarf Orchards from Monocular UAV Videos Using Global-Motion-Compensated DeepSORT

Dae-Hyun Lee, Longsheng Fu, Bryan Gilbert Murengami, Jun Chen, Yufei Dou, Lili SUN, Xiaopeng Yang, Shuangping Yang, Rui Li, Hongbing Meng
article en

Abstract

Fruit occlusion, dense clustering, platform vibration, and cross-frame re-identification errors reduce the reliability of video-stream yield estimation in high-density dwarf apple orchards. To address these limitations, this study developed a vision-based framework for apple detection, tracking, dynamic counting, and yield estimation. Apple detection was improved by embedding a Shuffle Attention (SA) module into YOLOv12n and introducing SlideLoss for hard-sample reweighting, yielding the YOLOv12n-SA-SlideLoss detector. Following detection, the DeepSORT tracking module was also improved by augmenting it with global motion compensation (GMC) and a dual-criteria association strategy based on motion consistency and appearance similarity. Yield estimation was formulated as a counting-based compensation model using dynamic fruit counts, mean fruit mass, and an occlusion compensation factor, and was compared with a pixel-geometric regression model based on detection boxes. Field experiments were conducted in a high-density dwarf ‘Fuji’ apple orchard in Aksu, Xinjiang, China. YOLOv12n-SA-SlideLoss achieved an [email protected] of 95.7%, exceeding the YOLOv12n baseline by 2.3%. The improved DeepSORT reached 87.4% multiple object tracking accuracy and 86.7% multiple object tracking precision, while reducing identification switches by 22.2% relative to the original algorithm. The dynamic counting framework increased average counting accuracy from 85.6% to 93.5%. For yield estimation, the counting-based compensation model achieved a mean prediction accuracy of 85.2%, an RMSE of 18.5 kg, and a mean bias of −17.2 kg across six validation samples (n = 6), each covering multiple trees. Overall, the results indicate that combining detection, motion-compensated tracking, and counting-based correction improves stability of single-view video-stream yield estimation in high-density dwarf apple orchards.

AgricultureVol. 16(19)
Chungnam National University (KR), Tarim University (CN), Ministry of Agriculture and Rural Affairs (CN), Northwest A&F University (CN)
Openalex Percentile: Top 14%
Smart Agriculture and AI
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.