Nutrient Calculation Accuracy, Expert-Rated Clinical Appropriateness and Safety Indicators, and Reproducibility of Large Language Model-Generated Diet Plans for Polyendocrine Metabolic Ovarian Syndrome: A Real-Case-Based Evaluation

Background/Objectives: This study evaluated the nutrient calculation accuracy, reproducibility, and expert-rated clinical appropriateness and safety indicators of three-day diet plans generated for women with polycystic ovary syndrome (PCOS), recently proposed to be termed polyendocrine metabolic ovarian syndrome (PMOS). Methods: Anonymized data from six women with heterogeneous PMOS profiles, purposively selected from a single clinic, were converted into standardized Turkish prompts and submitted to ChatGPT-4, Claude Sonnet 4.6, Gemini 2.5 Flash, and Mistral 2.0 through free web interfaces. Three independent generations per case–model combination yielded 216 daily menus. LLM-reported nutrient values were compared with BeBiS 9.0 calculations. Reproducibility was assessed using intraclass correlation coefficients (ICCs). Two experts independently evaluated the first generated plan for each case–model combination using a study-specific clinical appropriateness rubric and red-flag checklist. Results: No model achieved consistently low error across nutrients. Claude Sonnet 4.6 showed the lowest percentage errors for energy, carbohydrate, and fiber; ChatGPT-4 for protein; and Mistral 2.0 for fat. Gemini 2.5 Flash showed the largest percentage errors across all nutrients and a mean energy bias of +531.1 kcal. ChatGPT-4 showed comparatively higher plan-level reproducibility, while the remaining models generally showed poor reproducibility. Claude Sonnet 4.6 had the highest descriptive clinical appropriateness scores and Gemini 2.5 Flash the lowest, but no pairwise differences remained significant after Holm correction. Gemini 2.5 Flash generated the most confirmed study-specific red flags (n = 28), with all six cases meeting the criterion for the “masked low-energy” indicator. Conclusions: LLM-generated diet plans showed model-dependent calculation errors and limited reproducibility, and the present findings do not support their independent or unsupervised use as individualized clinical nutrition recommendations.

Authors

Institutions

Publication Details

Journal
Nutrients
Published
2026-09-28
DOI
https://doi.org/10.3390/nu18193205
Primary Topic
Ovarian function and disorders
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Nutrient Calculation Accuracy, Expert-Rated Clinical Appropriateness and Safety Indicators, and Reproducibility of Large Language Model-Generated Diet Plans for Polyendocrine Metabolic Ovarian Syndrome: A Real-Case-Based Evaluation

Pınar Ece Karakaş, Muazzez Garıpağaoğlu, Ayşe Betül Demirbaş, Ayşenur Emirhüseyinoğlu et al.
Nutrients
Ovarian function and disorders
article

Nutrient Calculation Accuracy, Expert-Rated Clinical Appropriateness and Safety Indicators, and Reproducibility of Large Language Model-Generated Diet Plans for Polyendocrine Metabolic Ovarian Syndrome: A Real-Case-Based Evaluation

Pınar Ece Karakaş, Muazzez Garıpağaoğlu, Ayşe Betül Demirbaş, Ayşenur Emirhüseyinoğlu, Gülen Ecem Kalkan
article en

Abstract

Background/Objectives: This study evaluated the nutrient calculation accuracy, reproducibility, and expert-rated clinical appropriateness and safety indicators of three-day diet plans generated for women with polycystic ovary syndrome (PCOS), recently proposed to be termed polyendocrine metabolic ovarian syndrome (PMOS). Methods: Anonymized data from six women with heterogeneous PMOS profiles, purposively selected from a single clinic, were converted into standardized Turkish prompts and submitted to ChatGPT-4, Claude Sonnet 4.6, Gemini 2.5 Flash, and Mistral 2.0 through free web interfaces. Three independent generations per case–model combination yielded 216 daily menus. LLM-reported nutrient values were compared with BeBiS 9.0 calculations. Reproducibility was assessed using intraclass correlation coefficients (ICCs). Two experts independently evaluated the first generated plan for each case–model combination using a study-specific clinical appropriateness rubric and red-flag checklist. Results: No model achieved consistently low error across nutrients. Claude Sonnet 4.6 showed the lowest percentage errors for energy, carbohydrate, and fiber; ChatGPT-4 for protein; and Mistral 2.0 for fat. Gemini 2.5 Flash showed the largest percentage errors across all nutrients and a mean energy bias of +531.1 kcal. ChatGPT-4 showed comparatively higher plan-level reproducibility, while the remaining models generally showed poor reproducibility. Claude Sonnet 4.6 had the highest descriptive clinical appropriateness scores and Gemini 2.5 Flash the lowest, but no pairwise differences remained significant after Holm correction. Gemini 2.5 Flash generated the most confirmed study-specific red flags (n = 28), with all six cases meeting the criterion for the “masked low-energy” indicator. Conclusions: LLM-generated diet plans showed model-dependent calculation errors and limited reproducibility, and the present findings do not support their independent or unsupervised use as individualized clinical nutrition recommendations.

NutrientsVol. 18(19)
Fenerbahçe University (TR), Atlas Üniversitesi
Zero hunger
Openalex Percentile: Top 9%
Ovarian function and disorders
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.