Synthetic Data for Data-Efficient Agricultural Computer Vision: Three Datasets and a Multi-Task Benchmark Across Real and Simulated Domains

The availability of large annotated datasets remains a major bottleneck for deploying computer vision systems in agricultural and agri-food applications, particularly for tasks requiring fine-grained annotations such as defect detection and instance segmentation. Synthetic data generation offers a scalable alternative, but its effectiveness in complex real-world scenarios remains insufficiently understood. In this work, the role of synthetic data is investigated across three representative tasks each posing unique challenges: apple detection in orchard environments, potato–stone classification in industrial sorting, and carrot crack detection for quality inspection. The real and synthetic data, as well as their annotations, are publicly available. The synthetic data is created using a set of tools for fast large-scale procedural scene generation with highly detailed natural assets. For each task, datasets are constructed combining real and synthetically generated images with pixel-level annotations. Training strategies are systematically evaluated using fully real, limited real, synthetic-only, and combined datasets. The results show that synthetic-only training leads to a clear performance gap on real-world data, with mAP50 decreasing from 0.866–0.891 for fully real training to 0.396–0.703 for synthetic-only training across the three use cases. However, combining synthetic data with only 10 real images recovers 75–98% of the gap between training on those 10 real images alone (mAP50 0.194–0.497) and full real training, reaching mAP50 values of 0.723–0.885. These findings highlight the potential of synthetic data for improving data efficiency and provide a reproducible benchmark for future research on sim-to-real transfer in agricultural computer vision.

Authors

Institutions

Publication Details

Journal
AgriEngineering
Published
2026-09-30
DOI
https://doi.org/10.3390/agriengineering8100414
Primary Topic
Smart Agriculture and AI
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Synthetic Data for Data-Efficient Agricultural Computer Vision: Three Datasets and a Multi-Task Benchmark Across Real and Simulated Domains

Mohammad Hasan Rahmani, Steven Moonen, Wenzhi Liao, Abdellatif Bey-Temsamani et al.
AgriEngineering
Smart Agriculture and AI
article

Synthetic Data for Data-Efficient Agricultural Computer Vision: Three Datasets and a Multi-Task Benchmark Across Real and Simulated Domains

Mohammad Hasan Rahmani, Steven Moonen, Wenzhi Liao, Abdellatif Bey-Temsamani, Jan Steckel, Wouter Jansen, Ayyoub Ahar, Nick Michiels
article en

Abstract

The availability of large annotated datasets remains a major bottleneck for deploying computer vision systems in agricultural and agri-food applications, particularly for tasks requiring fine-grained annotations such as defect detection and instance segmentation. Synthetic data generation offers a scalable alternative, but its effectiveness in complex real-world scenarios remains insufficiently understood. In this work, the role of synthetic data is investigated across three representative tasks each posing unique challenges: apple detection in orchard environments, potato–stone classification in industrial sorting, and carrot crack detection for quality inspection. The real and synthetic data, as well as their annotations, are publicly available. The synthetic data is created using a set of tools for fast large-scale procedural scene generation with highly detailed natural assets. For each task, datasets are constructed combining real and synthetically generated images with pixel-level annotations. Training strategies are systematically evaluated using fully real, limited real, synthetic-only, and combined datasets. The results show that synthetic-only training leads to a clear performance gap on real-world data, with mAP50 decreasing from 0.866–0.891 for fully real training to 0.396–0.703 for synthetic-only training across the three use cases. However, combining synthetic data with only 10 real images recovers 75–98% of the gap between training on those 10 real images alone (mAP50 0.194–0.497) and full real training, reaching mAP50 values of 0.723–0.885. These findings highlight the potential of synthetic data for improving data efficiency and provide a reproducible benchmark for future research on sim-to-real transfer in agricultural computer vision.

AgriEngineeringVol. 8(10)
University of Antwerp (BE), Flanders Make (Belgium) (BE), Hasselt University (BE)
Zero hunger
Openalex Percentile: Top 14%
Smart Agriculture and AI
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.