Internal Consistency Does Not Validate an LLM-Assisted Nutrition Pipeline: A Reference-Standard-Free Forensic Reliability Audit of 13,073 Free-Text Meal Records

Background/Objectives: Large language model (LLM) pipelines increasingly convert free-text meal descriptions into nutrient values, yet reference-standard validation is rarely feasible in production. We asked whether the controls such pipelines conventionally pass suffice and built a reference-standard-free battery any pipeline can run on its output. Methods: We audited 13,073 meals logged by 513 adults through a WhatsApp conversational agent in rural Colombia (9 February 2026 to 6 August 2026). The pipeline is a two-stage cascade of commercial LLMs from different providers: a vision model renders photographs as text, while a text model maps these to nutrients. Six pre-specified tests were applied, including test–retest ICC(1,1) over recurring strings, stratified by modality because 48.3% of records carried a photograph. Results: The pipeline passed every conventional control (Atwater-reconstructed energy r = 0.9951; median relative error 0.0000) but failed reproducibility. Across 191 recurring strings (944 meals), energy ICC point estimates were 0.7061 (95% CI 0.5419–0.8217) text only and 0.5258 (0.3778–0.6513) with photographs; none reached 0.75. Modality is a design control, not an established effect, and the coefficient depends on which inputs recur, so it is not intrinsic. Under the strictest identity available—byte-identical text, no photograph—ICC was 0.6944. Intermediate-text inspection exposed a failure invisible at the output: 1.44% of photograph-accompanied records carried a vision-stage artefact string versus 0.01% of text-only records (ICC 0.0911). Conclusions: Internal consistency does not validate an LLM-assisted nutrition pipeline. This design measures instability under recurring inputs without isolating its cause. Prospective test–retest reproducibility is a cheap, necessary acceptance criterion, and in a cascade, the intermediate representation must also be inspected.

Authors

Institutions

Publication Details

Journal
Nutrients
Published
2026-09-30
DOI
https://doi.org/10.3390/nu18193233
Primary Topic
Child Nutrition and Water Access
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Internal Consistency Does Not Validate an LLM-Assisted Nutrition Pipeline: A Reference-Standard-Free Forensic Reliability Audit of 13,073 Free-Text Meal Records

Edwin Rivas Trujillo, Juan Gabriel Castañeda Polanco, Jhony F. López
Nutrients
Child Nutrition and Water Access
article

Internal Consistency Does Not Validate an LLM-Assisted Nutrition Pipeline: A Reference-Standard-Free Forensic Reliability Audit of 13,073 Free-Text Meal Records

Edwin Rivas Trujillo, Juan Gabriel Castañeda Polanco, Jhony F. López
article en

Abstract

Background/Objectives: Large language model (LLM) pipelines increasingly convert free-text meal descriptions into nutrient values, yet reference-standard validation is rarely feasible in production. We asked whether the controls such pipelines conventionally pass suffice and built a reference-standard-free battery any pipeline can run on its output. Methods: We audited 13,073 meals logged by 513 adults through a WhatsApp conversational agent in rural Colombia (9 February 2026 to 6 August 2026). The pipeline is a two-stage cascade of commercial LLMs from different providers: a vision model renders photographs as text, while a text model maps these to nutrients. Six pre-specified tests were applied, including test–retest ICC(1,1) over recurring strings, stratified by modality because 48.3% of records carried a photograph. Results: The pipeline passed every conventional control (Atwater-reconstructed energy r = 0.9951; median relative error 0.0000) but failed reproducibility. Across 191 recurring strings (944 meals), energy ICC point estimates were 0.7061 (95% CI 0.5419–0.8217) text only and 0.5258 (0.3778–0.6513) with photographs; none reached 0.75. Modality is a design control, not an established effect, and the coefficient depends on which inputs recur, so it is not intrinsic. Under the strictest identity available—byte-identical text, no photograph—ICC was 0.6944. Intermediate-text inspection exposed a failure invisible at the output: 1.44% of photograph-accompanied records carried a vision-stage artefact string versus 0.01% of text-only records (ICC 0.0911). Conclusions: Internal consistency does not validate an LLM-assisted nutrition pipeline. This design measures instability under recurring inputs without isolating its cause. Prospective test–retest reproducibility is a cheap, necessary acceptance criterion, and in a cascade, the intermediate representation must also be inspected.

NutrientsVol. 18(19)
University of Tolima (CO), Universidad Distrital Francisco José de Caldas (CO), Corporación Universitaria Minuto de Dios (CO)
Zero hunger
Openalex Percentile: Top 13%
Child Nutrition and Water Access
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.