A Deployed Vision–Language Decision-Support Platform for Tomato Harvest: Identity-Gated Grading, Uncertainty-Propagating Yield Estimation and Weather-Conditioned Advisories
Most crop-vision research stops at detection: a bounding box and a class label. A grower’s decision needs more — how much marketable weight is on the plant, how much of it is already lost, and whether to pick today or wait four days. Vision–language models (VLMs) make the whole chain buildable without a labelled dataset, and that changes which engineering problems dominate. We describe and measure a deployed six-language mobile platform, built for a horticultural research institute, that grades tomatoes against a six-stage ripeness taxonomy, estimates marketable yield with propagated uncertainty, assesses damage and projects weather-driven loss, and issues advisories from a deterministic rule layer conditioned on a three-provider weather forecast. The system is 37,626 lines across 21 endpoints, passes 851 automated tests, and runs on real handsets. Our contributions are an architecture and four measurements. First, identity gating: a prompt that asserted the premise it was meant to test reported tomatoes in 19 of 20 photographs confirmed to contain none, emitting 92 false boxes; restructuring the prompt to identify before grading, and adding identity fields to the constrained-decoding schema, gave 0 of 19 and 0 false boxes on the same pinned model with recall unchanged, and the deployed primary engine’s identity precision is 28/28 on a labelled negative set. Second, run-to-run reliability: 35 field photographs analysed three times each (105 runs, 816 fruit) agree on ripeness 93.1% of the time over 1,448 matched fruit (κ = 0.902), every confusion between adjacent stages, with breaker — the stage that decides wait-versus-harvest — least stable at F1 = 0.730; because two disagreeing runs cannot both be right, this yields a derived accuracy ceiling of 97.9% obtainable without labels. Third, intermittent detection collapse and its cause: 6.7% of runs returned approximately one fruit where siblings found 11–18, and the cause is neither the model nor the reasoning setting but the same under-specified prompt (26.9% versus 0.0% severe under-counts, one-sided Fisher exact p = 0.0037). Fourth, an estimation architecture in which every downstream number carries the uncertainty of its weakest input: geometry corroborates but never overrides an absolute size judgement, markerless yield confidence is capped at 0.65 by construction, reference evapotranspiration returns no value rather than zero when inputs are insufficient, and dependent advisories are skipped rather than computed from a placeholder. We state plainly what is not measured: no stage-labelled ground truth and no weighed-harvest calibration set existed, so ripeness classification accuracy and absolute yield error are not quantified, and the yield figures are defensible as relative comparisons rather than validated absolute masses.
Authors
- Manjunath Suresh
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-25
- DOI
- https://doi.org/10.5281/zenodo.22964165
- Primary Topic
- Smart Agriculture and AI
- Type
- article
- Field-Weighted Citation Impact
- 0.00