Nutritional Accuracy and Clinical Applicability of Large Language Model-Generated Therapeutic Diet Plans for Phenylketonuria: A Comparative Simulation Study

Background: Large language models (LLMs) are increasingly used to generate nutrition-related recommendations, but their reliability for therapeutic diets requiring precise nutrient control remains uncertain. This study evaluated AI-generated one-day diet plans for standardized simulated cases of classical phenylketonuria (PKU). Methods: Eight standardized simulated cases representing infancy, adolescence, adulthood, and maternal PKU were evaluated. Between 8 and 10 February 2026, ChatGPT-5.2, Claude 4.5, and Gemini 3 were each used through their standard consumer-facing chat interfaces to generate one diet plan per case under two prompting conditions, yielding 48 plans. Nutrient composition was independently recalculated using BİAYS, and three blinded pediatric metabolic dietitians evaluated the plans using a structured 25-item checklist. Results: Across all 48 plans, phenylalanine provision calculated using BİAYS was below the case-specific prescription in 36 plans (75.0%). After Holm correction, none of the 12 paired AI-reported versus BİAYS-calculated nutrient comparisons remained statistically significant, while plan-level error metrics demonstrated substantial variability for several nutrients. Overall expert scores did not differ among models, whereas clinically specified structured prompts received higher expert ratings than limited-information zero-shot prompts. Inter-rater reliability was modest at the diet-plan level. Conclusions: AI-generated PKU diet plans may incorporate core dietary principles, but numerical plausibility does not ensure agreement with independent nutrient calculations or adherence to individualized clinical prescriptions. Independent nutrient verification and specialist metabolic-dietitian review remain necessary. Because only one output was generated for each case–model–prompt combination, the findings characterize the 48 evaluated outputs rather than within-model reproducibility.

Authors

Institutions

Publication Details

Journal
Nutrients
Published
2026-09-25
DOI
https://doi.org/10.3390/nu18193158
Primary Topic
Metabolism and Genetic Disorders
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Nutritional Accuracy and Clinical Applicability of Large Language Model-Generated Therapeutic Diet Plans for Phenylketonuria: A Comparative Simulation Study

Banu Süzen, Nuriye Ece Mintaş, Volkan Hancı, Bahar Kulu et al.
Nutrients
Metabolism and Genetic Disorders
article

Nutritional Accuracy and Clinical Applicability of Large Language Model-Generated Therapeutic Diet Plans for Phenylketonuria: A Comparative Simulation Study

Banu Süzen, Nuriye Ece Mintaş, Volkan Hancı, Bahar Kulu, Sibel Burçak Şahin Uyar, Nur Arslan, Fulya Durğun, Neslihan Buse Tatlıdil, Esra Özcan
article en

Abstract

Background: Large language models (LLMs) are increasingly used to generate nutrition-related recommendations, but their reliability for therapeutic diets requiring precise nutrient control remains uncertain. This study evaluated AI-generated one-day diet plans for standardized simulated cases of classical phenylketonuria (PKU). Methods: Eight standardized simulated cases representing infancy, adolescence, adulthood, and maternal PKU were evaluated. Between 8 and 10 February 2026, ChatGPT-5.2, Claude 4.5, and Gemini 3 were each used through their standard consumer-facing chat interfaces to generate one diet plan per case under two prompting conditions, yielding 48 plans. Nutrient composition was independently recalculated using BİAYS, and three blinded pediatric metabolic dietitians evaluated the plans using a structured 25-item checklist. Results: Across all 48 plans, phenylalanine provision calculated using BİAYS was below the case-specific prescription in 36 plans (75.0%). After Holm correction, none of the 12 paired AI-reported versus BİAYS-calculated nutrient comparisons remained statistically significant, while plan-level error metrics demonstrated substantial variability for several nutrients. Overall expert scores did not differ among models, whereas clinically specified structured prompts received higher expert ratings than limited-information zero-shot prompts. Inter-rater reliability was modest at the diet-plan level. Conclusions: AI-generated PKU diet plans may incorporate core dietary principles, but numerical plausibility does not ensure agreement with independent nutrient calculations or adherence to individualized clinical prescriptions. Independent nutrient verification and specialist metabolic-dietitian review remain necessary. Because only one output was generated for each case–model–prompt combination, the findings characterize the 48 evaluated outputs rather than within-model reproducibility.

NutrientsVol. 18(19)
Izmir University (TR), Dokuz Eylül University (TR), Cappadocia University (TR), Sağlık Bilimleri Üniversitesi (TR), Izmir Tepecik Eğitim ve Araştırma Hastanesi (TR), Muğla University (TR)
Zero hunger
Openalex Percentile: Top 15%
Metabolism and Genetic Disorders
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.