Domain-Dependent Performance of Human Experts and AI Systems in Pediatric Menu Evaluation

Background/Objectives: Artificial intelligence (AI) tools are increasingly used for dietary assessment, but their reliability in pediatric nutrition remains uncertain. This study compared AI systems and human evaluators in pediatric menu assessment against expert-defined reference ratings. Methods: This observational cross-sectional study used anonymized dietary and clinical information from three healthy pediatric cases. A multidisciplinary panel of pediatric nutrition specialists established standard evaluations. Assessments were completed by 84 AI evaluations, 116 nutrition specialists, 56 pediatric-focused physicians, and 88 physicians from other specialties. Outcomes included absolute error in total energy estimation, deviations in portion and qualitative ratings, standardized scores, and binary adequacy accuracy. Groups were compared using ANOVA or Kruskal–Wallis tests with post hoc analyses. Results: A total of 344 assessments were included. Energy-estimation error differed between groups (Kruskal–Wallis = 12.85, p = 0.0016), but this analysis was based on a limited and highly unbalanced subset of 77 assessments. AI had the largest mean absolute error (260.3 ± 194.3 kcal), followed by nutrition specialists (148.4 ± 81.5 kcal); physicians with other specialties had the lowest error (60.4 ± 53.3 kcal); these subgroup comparisons should be interpreted cautiously because of the small and unequal numbers of available energy estimates. AI showed greater portion-rating deviation (0.488; 95% CI: 0.38–0.60) than nutrition specialists (0.310; 95% CI: 0.20–0.42; p = 0.0019, q = 0.0059) and pediatric-focused physicians (0.286; 95% CI: 0.16–0.41; p = 0.0181, q = 0.019). Conversely, AI showed the lowest deviations for variety (0.37 ± 0.49) and processing (0.33 ± 0.47), compared with 1.15–1.36 and 1.22–1.57, respectively, among human groups (both p = 0.0001). Conclusions: AI may support structured pediatric menu screening for descriptive qualitative features; however, lower precision for energy and portion assessment supports, under these specific conditions, its use as an adjunct, for qualified nutrition professionals.

Authors

Institutions

Publication Details

Journal
Nutrients
Published
2026-09-21
DOI
https://doi.org/10.3390/nu18183106
Primary Topic
Nutrition, Genetics, and Disease
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Domain-Dependent Performance of Human Experts and AI Systems in Pediatric Menu Evaluation

Monica Tarcea, Roxana Maria Martin-Hadmaș, George Mihăiță Gavra, Ștefan Adrian Martin et al.
Nutrients
Nutrition, Genetics, and Disease
article

Domain-Dependent Performance of Human Experts and AI Systems in Pediatric Menu Evaluation

Monica Tarcea, Roxana Maria Martin-Hadmaș, George Mihăiță Gavra, Ștefan Adrian Martin, Rebeca Sovea, Adriana Neghirlă, Diana Pol
article en

Abstract

Background/Objectives: Artificial intelligence (AI) tools are increasingly used for dietary assessment, but their reliability in pediatric nutrition remains uncertain. This study compared AI systems and human evaluators in pediatric menu assessment against expert-defined reference ratings. Methods: This observational cross-sectional study used anonymized dietary and clinical information from three healthy pediatric cases. A multidisciplinary panel of pediatric nutrition specialists established standard evaluations. Assessments were completed by 84 AI evaluations, 116 nutrition specialists, 56 pediatric-focused physicians, and 88 physicians from other specialties. Outcomes included absolute error in total energy estimation, deviations in portion and qualitative ratings, standardized scores, and binary adequacy accuracy. Groups were compared using ANOVA or Kruskal–Wallis tests with post hoc analyses. Results: A total of 344 assessments were included. Energy-estimation error differed between groups (Kruskal–Wallis = 12.85, p = 0.0016), but this analysis was based on a limited and highly unbalanced subset of 77 assessments. AI had the largest mean absolute error (260.3 ± 194.3 kcal), followed by nutrition specialists (148.4 ± 81.5 kcal); physicians with other specialties had the lowest error (60.4 ± 53.3 kcal); these subgroup comparisons should be interpreted cautiously because of the small and unequal numbers of available energy estimates. AI showed greater portion-rating deviation (0.488; 95% CI: 0.38–0.60) than nutrition specialists (0.310; 95% CI: 0.20–0.42; p = 0.0019, q = 0.0059) and pediatric-focused physicians (0.286; 95% CI: 0.16–0.41; p = 0.0181, q = 0.019). Conversely, AI showed the lowest deviations for variety (0.37 ± 0.49) and processing (0.33 ± 0.47), compared with 1.15–1.36 and 1.22–1.57, respectively, among human groups (both p = 0.0001). Conclusions: AI may support structured pediatric menu screening for descriptive qualitative features; however, lower precision for energy and portion assessment supports, under these specific conditions, its use as an adjunct, for qualified nutrition professionals.

NutrientsVol. 18(18)
Universitatea de Medicină, Farmacie, Științe și Tehnologie „George Emil Palade” din Târgu Mureș (RO), Spitalul Clinic Judetean de Urgenta Târgu Mureş (RO), University of Arts from Târgu-Mureș (RO)
Zero hunger
Openalex Percentile: Top 11%
Nutrition, Genetics, and Disease
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.