Flatfoot and artificial intelligence

Flatfoot (pes planus) is common in children and adults. Although often asymptomatic, it may alter lower-limb biomechanics and contribute to pain or injury. Patients and caregivers increasingly use artificial intelligence chatbots such as chat generative pretrained transformer (ChatGPT) and Google Gemini for medical information, yet their responses on flatfoot remain insufficiently studied. This study compared the factual accuracy, added value, omissions, and readability of responses generated by ChatGPT and Google Gemini to standardized patient-oriented questions. In this cross-sectional comparative study, 15 standardized patient-oriented questions were developed from clinical guidelines, peer-reviewed literature, and educational resources from established orthopedic societies, then refined by 2 orthopedic surgeons. Each question was submitted to ChatGPT (OpenAI) and Google Gemini (Google LLC) in July 2025 using identical wording in independent chat sessions. Responses were anonymized and independently rated by 2 board-certified orthopedic surgeons for factual accuracy, added value, and omissions using an investigator-developed 5-point rubric. Discrepant assessments were reviewed by a 3rd senior orthopedic surgeon. Primary analyses used mean scores of the 2 reviewers. Readability was assessed using the Flesch–Kincaid grade level and Flesch reading ease score. Paired differences were analyzed using the Wilcoxon signed-rank test, with P < .05 considered significant. Both models produced generally accurate responses, with mean factual accuracy scores above 4/5. Gemini scored higher than ChatGPT for factual accuracy (4.7 ± 0.2 vs 4.2 ± 0.3, P = .01), added value (4.6 ± 0.3 vs 4.0 ± 0.4, P = .02), and omissions (4.8 ± 0.2 vs 3.9 ± 0.4, P < .001), with higher scores indicating fewer clinically relevant omissions. Gemini also generated longer responses (165 ± 20 vs 120 ± 15 words, P < .001) and showed better readability, with lower Flesch–Kincaid grade level (9.2 ± 1.1 vs 11.8 ± 1.4, P < .001) and higher Flesch reading ease score (59.4 ± 6.2 vs 42.5 ± 5.3, P < .001). Both models generated generally accurate answers to standardized flatfoot questions. Gemini performed better across evaluation domains and readability measures. As these findings are time-specific, repeated assessments and direct patient and caregiver evaluations are needed before broader patient-facing use can be recommended.

Authors

Institutions

Publication Details

Journal
Medicine
Published
2026-09-25
DOI
https://doi.org/10.1097/md.0000000000050839
Primary Topic
Lower Extremity Biomechanics and Pathologies
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Flatfoot and artificial intelligence

Ruhat Ünlü, Hasan Emirhan Usta
Medicine
Lower Extremity Biomechanics and Pathologies
article

Flatfoot and artificial intelligence

Ruhat Ünlü, Hasan Emirhan Usta
article en

Abstract

Flatfoot (pes planus) is common in children and adults. Although often asymptomatic, it may alter lower-limb biomechanics and contribute to pain or injury. Patients and caregivers increasingly use artificial intelligence chatbots such as chat generative pretrained transformer (ChatGPT) and Google Gemini for medical information, yet their responses on flatfoot remain insufficiently studied. This study compared the factual accuracy, added value, omissions, and readability of responses generated by ChatGPT and Google Gemini to standardized patient-oriented questions. In this cross-sectional comparative study, 15 standardized patient-oriented questions were developed from clinical guidelines, peer-reviewed literature, and educational resources from established orthopedic societies, then refined by 2 orthopedic surgeons. Each question was submitted to ChatGPT (OpenAI) and Google Gemini (Google LLC) in July 2025 using identical wording in independent chat sessions. Responses were anonymized and independently rated by 2 board-certified orthopedic surgeons for factual accuracy, added value, and omissions using an investigator-developed 5-point rubric. Discrepant assessments were reviewed by a 3rd senior orthopedic surgeon. Primary analyses used mean scores of the 2 reviewers. Readability was assessed using the Flesch–Kincaid grade level and Flesch reading ease score. Paired differences were analyzed using the Wilcoxon signed-rank test, with P < .05 considered significant. Both models produced generally accurate responses, with mean factual accuracy scores above 4/5. Gemini scored higher than ChatGPT for factual accuracy (4.7 ± 0.2 vs 4.2 ± 0.3, P = .01), added value (4.6 ± 0.3 vs 4.0 ± 0.4, P = .02), and omissions (4.8 ± 0.2 vs 3.9 ± 0.4, P < .001), with higher scores indicating fewer clinically relevant omissions. Gemini also generated longer responses (165 ± 20 vs 120 ± 15 words, P < .001) and showed better readability, with lower Flesch–Kincaid grade level (9.2 ± 1.1 vs 11.8 ± 1.4, P < .001) and higher Flesch reading ease score (59.4 ± 6.2 vs 42.5 ± 5.3, P < .001). Both models generated generally accurate answers to standardized flatfoot questions. Gemini performed better across evaluation domains and readability measures. As these findings are time-specific, repeated assessments and direct patient and caregiver evaluations are needed before broader patient-facing use can be recommended.

MedicineVol. 105(39)
Istanbul University (TR)
Quality Education
Openalex Percentile: Top 21%
Lower Extremity Biomechanics and Pathologies
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.