Reliability Interventions for Vision-Language Extraction from Real Nutrition Labels: A Replicated Held-Out Evaluation
Hosted vision-language models (VLMs) are used for key-value extraction from document photographs, often wrapped in client-side reliability interventions. We test whether image preprocessing, multi-crop inference and schema validation with one corrective retry improve numeric extraction from real Nutrition Facts labels, and whether a synthetic development benchmark predicts their effect. One hosted VLM (Claude Sonnet 5) processed a frozen held-out set of 100 Open Food Facts photographs under four input conditions, three times each. We score the 453 of 600 candidate numeric fields whose image-level values could be independently verified by OCR; the reported accuracy therefore characterizes this OCR-verifiable subset rather than all label fields. Accuracy was near ceiling in every condition (product-level mean 0.997–0.998), and all six paired product-level bootstrap intervals for pairwise differences included zero. Multi-crop raised cost per image by 84% with no net stable rescue. Schema validation found one retry opportunity in 300 responses; the retry restored validity without changing any scored value. The intervention ranking and the accuracy, stability and cost estimates from our synthetic development benchmark did not reproduce on the real held-out set. On this image distribution, the failure modes these interventions target were rare, leaving little measurable headroom over raw input; no intervention produced a detectable improvement, which is not evidence of equivalence. Findings are limited to this model and label type.
Authors
- Aditya Aralimara
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-27
- DOI
- https://doi.org/10.5281/zenodo.22982305
- Primary Topic
- Biomedical Text Mining and Ontologies
- Type
- preprint