Reliability Interventions for Vision-Language Extraction from Real Nutrition Labels: A Replicated Held-Out Evaluation

Hosted vision-language models (VLMs) are used for key-value extraction from document photographs, often wrapped in client-side reliability interventions. We test whether image preprocessing, multi-crop inference and schema validation with one corrective retry improve numeric extraction from real Nutrition Facts labels, and whether a synthetic development benchmark predicts their effect. One hosted VLM (Claude Sonnet 5) processed a frozen held-out set of 100 Open Food Facts photographs under four input conditions, three times each. We score the 453 of 600 candidate numeric fields whose image-level values could be independently verified by OCR; the reported accuracy therefore characterizes this OCR-verifiable subset rather than all label fields. Accuracy was near ceiling in every condition (product-level mean 0.997–0.998), and all six paired product-level bootstrap intervals for pairwise differences included zero. Multi-crop raised cost per image by 84% with no net stable rescue. Schema validation found one retry opportunity in 300 responses; the retry restored validity without changing any scored value. The intervention ranking and the accuracy, stability and cost estimates from our synthetic development benchmark did not reproduce on the real held-out set. On this image distribution, the failure modes these interventions target were rare, leaving little measurable headroom over raw input; no intervention produced a detectable improvement, which is not evidence of equivalence. Findings are limited to this model and label type.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-27
DOI
https://doi.org/10.5281/zenodo.22982306
Primary Topic
Biomedical Text Mining and Ontologies
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Reliability Interventions for Vision-Language Extraction from Real Nutrition Labels: A Replicated Held-Out Evaluation

Aditya Aralimara
Zenodo (CERN European Organization for Nuclear Research)
Biomedical Text Mining and Ontologies
preprint

Reliability Interventions for Vision-Language Extraction from Real Nutrition Labels: A Replicated Held-Out Evaluation

Aditya Aralimara
preprint en

Abstract

Hosted vision-language models (VLMs) are used for key-value extraction from document photographs, often wrapped in client-side reliability interventions. We test whether image preprocessing, multi-crop inference and schema validation with one corrective retry improve numeric extraction from real Nutrition Facts labels, and whether a synthetic development benchmark predicts their effect. One hosted VLM (Claude Sonnet 5) processed a frozen held-out set of 100 Open Food Facts photographs under four input conditions, three times each. We score the 453 of 600 candidate numeric fields whose image-level values could be independently verified by OCR; the reported accuracy therefore characterizes this OCR-verifiable subset rather than all label fields. Accuracy was near ceiling in every condition (product-level mean 0.997–0.998), and all six paired product-level bootstrap intervals for pairwise differences included zero. Multi-crop raised cost per image by 84% with no net stable rescue. Schema validation found one retry opportunity in 300 responses; the retry restored validity without changing any scored value. The intervention ranking and the accuracy, stability and cost estimates from our synthetic development benchmark did not reproduce on the real held-out set. On this image distribution, the failure modes these interventions target were rare, leaving little measurable headroom over raw input; no intervention produced a detectable improvement, which is not evidence of equivalence. Findings are limited to this model and label type.

Zenodo (CERN European Organization for Nuclear Research)
Zero hunger
Biomedical Text Mining and Ontologies
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.