Nutritional Accuracy and Clinical Applicability of Large Language Model-Generated Therapeutic Diet Plans for Phenylketonuria: A Comparative Simulation Study
Background: Large language models (LLMs) are increasingly used to generate nutrition-related recommendations, but their reliability for therapeutic diets requiring precise nutrient control remains uncertain. This study evaluated AI-generated one-day diet plans for standardized simulated cases of classical phenylketonuria (PKU). Methods: Eight standardized simulated cases representing infancy, adolescence, adulthood, and maternal PKU were evaluated. Between 8 and 10 February 2026, ChatGPT-5.2, Claude 4.5, and Gemini 3 were each used through their standard consumer-facing chat interfaces to generate one diet plan per case under two prompting conditions, yielding 48 plans. Nutrient composition was independently recalculated using BİAYS, and three blinded pediatric metabolic dietitians evaluated the plans using a structured 25-item checklist. Results: Across all 48 plans, phenylalanine provision calculated using BİAYS was below the case-specific prescription in 36 plans (75.0%). After Holm correction, none of the 12 paired AI-reported versus BİAYS-calculated nutrient comparisons remained statistically significant, while plan-level error metrics demonstrated substantial variability for several nutrients. Overall expert scores did not differ among models, whereas clinically specified structured prompts received higher expert ratings than limited-information zero-shot prompts. Inter-rater reliability was modest at the diet-plan level. Conclusions: AI-generated PKU diet plans may incorporate core dietary principles, but numerical plausibility does not ensure agreement with independent nutrient calculations or adherence to individualized clinical prescriptions. Independent nutrient verification and specialist metabolic-dietitian review remain necessary. Because only one output was generated for each case–model–prompt combination, the findings characterize the 48 evaluated outputs rather than within-model reproducibility.
Authors
- Banu Süzen (ORCID: https://orcid.org/0000-0002-5975-5868)
- Nuriye Ece Mintaş (ORCID: https://orcid.org/0000-0002-0033-6861)
- Volkan Hancı (ORCID: https://orcid.org/0000-0002-2227-194X)
- Bahar Kulu (ORCID: https://orcid.org/0000-0003-2147-9316)
- Sibel Burçak Şahin Uyar
- Nur Arslan
- Fulya Durğun (ORCID: https://orcid.org/0009-0005-1365-2475)
- Neslihan Buse Tatlıdil (ORCID: https://orcid.org/0009-0002-0059-178X)
- Esra Özcan (ORCID: https://orcid.org/0009-0009-9393-143X)
Institutions
- Izmir University (TR)
- Dokuz Eylül University (TR)
- Cappadocia University (TR)
- Sağlık Bilimleri Üniversitesi (TR)
- Izmir Tepecik Eğitim ve Araştırma Hastanesi (TR)
- Muğla University (TR)
Publication Details
- Journal
- Nutrients
- Published
- 2026-09-25
- DOI
- https://doi.org/10.3390/nu18193158
- Primary Topic
- Metabolism and Genetic Disorders
- Type
- article
- Field-Weighted Citation Impact
- 0.00