Evidence integrity and review utility in crop-image decision support: an audit-to-review framework for precision agriculture

Abstract Background Agricultural artificial intelligence benchmarks can overstate decision reliability when multiple files represent the same biological evidence or when split provenance is unclear. We evaluated a classifier-agnostic audit-to-review framework that fixes evidence identity and role provenance before model comparison, declares aggregation estimands, and links predictive uncertainty to review workload. A public wheat-leaf corpus of 7,595 files was hashed to 5,118 unique contents; eligibility criteria defined a 915-content five-class task. Archived file-level benchmarks were separated from three frozen near-duplicate-grouped P2 partition configurations, and semantic, texture and fusion representations were evaluated under partition-weighted and one-content-one-weight estimands. Split conformal prediction was assessed by empirical coverage, review rate, error capture, review yield and dimensionless review utility. Results Macro-F1 was 0.962 under the legacy P0 benchmark, 0.805 ± 0.030 for semantic-only P2 and 0.819 ± 0.029 for direct fusion. The descriptive P0-P2 workflow-sensitivity gap was approximately 15.7% points, versus a 1.4-point partition-weighted fusion increment (95% bootstrap interval − 0.006 to 0.036). The leading fusion-related variants were numerically close, so no inferential ranking is claimed. At α = 0.10, global split conformal prediction achieved 0.896 empirical coverage while reviewing 20.8% of images and capturing 51.4% of top-1 errors; class-conditional Mondrian calibration achieved 0.923 coverage while reviewing 29.7% and capturing 64.8% of errors. Conclusions Defining the evidence unit is part of the statistical specification of a crop-image benchmark, not merely a data-cleaning step. In this corpus, benchmark interpretation was much more sensitive to evidence-workflow specification than to the evaluated representation refinement. The framework supports auditable internal evaluation and review triage but does not establish field, farm-level or intervention performance.

Authors

Institutions

Publication Details

Journal
Bulletin of the National Research Centre/Bulletin of the National Research Center
Published
2026-09-14
DOI
https://doi.org/10.1186/s42269-026-01492-x
Primary Topic
Smart Agriculture and AI
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Evidence integrity and review utility in crop-image decision support: an audit-to-review framework for precision agriculture

A.Dzh. Kartanova, Bojian Guo, Xin Li
Bulletin of the National Research Centre/Bulletin of the National Research Center
Smart Agriculture and AI
article

Evidence integrity and review utility in crop-image decision support: an audit-to-review framework for precision agriculture

A.Dzh. Kartanova, Bojian Guo, Xin Li
article en

Abstract

Abstract Background Agricultural artificial intelligence benchmarks can overstate decision reliability when multiple files represent the same biological evidence or when split provenance is unclear. We evaluated a classifier-agnostic audit-to-review framework that fixes evidence identity and role provenance before model comparison, declares aggregation estimands, and links predictive uncertainty to review workload. A public wheat-leaf corpus of 7,595 files was hashed to 5,118 unique contents; eligibility criteria defined a 915-content five-class task. Archived file-level benchmarks were separated from three frozen near-duplicate-grouped P2 partition configurations, and semantic, texture and fusion representations were evaluated under partition-weighted and one-content-one-weight estimands. Split conformal prediction was assessed by empirical coverage, review rate, error capture, review yield and dimensionless review utility. Results Macro-F1 was 0.962 under the legacy P0 benchmark, 0.805 ± 0.030 for semantic-only P2 and 0.819 ± 0.029 for direct fusion. The descriptive P0-P2 workflow-sensitivity gap was approximately 15.7% points, versus a 1.4-point partition-weighted fusion increment (95% bootstrap interval − 0.006 to 0.036). The leading fusion-related variants were numerically close, so no inferential ranking is claimed. At α = 0.10, global split conformal prediction achieved 0.896 empirical coverage while reviewing 20.8% of images and capturing 51.4% of top-1 errors; class-conditional Mondrian calibration achieved 0.923 coverage while reviewing 29.7% and capturing 64.8% of errors. Conclusions Defining the evidence unit is part of the statistical specification of a crop-image benchmark, not merely a data-cleaning step. In this corpus, benchmark interpretation was much more sensitive to evidence-workflow specification than to the evaluated representation refinement. The framework supports auditable internal evaluation and review triage but does not establish field, farm-level or intervention performance.

Bulletin of the National Research Centre/Bulletin of the National Research CenterVol. 50(1)
National Research University Higher School of Economics (RU), Kyrgyz State Technical University (KG)
Zero hunger
Openalex Percentile: Top 13%
Smart Agriculture and AI
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.