Comparative Performance of Large Language Models in Urodynamic Trace Interpretation Using an Adapted Expert‐Led Evaluation of Generative AI Competence and Excellence Framework

OBJECTIVES: Interpretation of urodynamic studies is complex and subject to substantial inter-observer variability, even among experienced clinicians. While artificial intelligence has shown promise in urology, evidence regarding the performance of large language models (LLMs) in urodynamic trace interpretation remains limited. This study aimed to systematically compare the performance of multiple contemporary LLMs in interpreting standard urodynamic traces using an adapted ELEGANCE framework. METHODS: In this retrospective, single-centre comparative study, 119 anonymised urodynamic machine printouts retrieved from the institutional patient archive were independently interpreted by five LLM-based platforms (ChatGPT-5.2, Gemini 3, Perplexity AI Pro, DeepSeek-V3.2, and Grok 4.1). Case-specific expert reference interpretations were produced by one of three urologists, whereas all LLM-generated interpretations were independently evaluated by all three urologists. LLM-generated interpretations were assessed using an adapted version of the "Expert-Led Evaluation of Generative AI Competence and Excellence" (ELEGANCE) questionnaire, which evaluates relevance, completeness, applicability, structure, language/terminology, satisfaction, and hallucination. Between-model comparisons were performed using Friedman and post-hoc Wilcoxon signed-rank tests. RESULTS: [4] = 12.67, p = 0.013). Gemini achieved the highest mean score (25.77 ± 3.74), followed by Perplexity, DeepSeek, and ChatGPT-5.2, while Grok demonstrated significantly lower performance. Gemini consistently outperformed other models in clinically oriented domains, including relevance, completeness, and applicability (all p < 0.05). No hallucinations were identified in any model output. Inter-rater reliability among the three expert reviewers was good (average-measures ICC = 0.898, 95% CI 0.853-0.929; p < 0.001). CONCLUSION: Contemporary LLMs can generate structured interpretations of urodynamic traces, with meaningful performance differences across the evaluated platforms. While strengths in structure and terminology are evident, limitations in clinical reasoning persist. LLMs may serve as valuable assistive tools for standardisation and quality assurance in urodynamics but are not yet suitable for autonomous clinical decision-making.

Authors

Institutions

Publication Details

Journal
Neurourology and Urodynamics
Published
2026-10-05
DOI
https://doi.org/10.1002/nau.70470
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Comparative Performance of Large Language Models in Urodynamic Trace Interpretation Using an Adapted Expert‐Led Evaluation of Generative AI Competence and Excellence Framework

Muhammet Güzelsoy, Anıl Erkan, Oğuzhan AKPINAR, Alper Keskin
Neurourology and Urodynamics
Artificial Intelligence in Healthcare and Education
article

Comparative Performance of Large Language Models in Urodynamic Trace Interpretation Using an Adapted Expert‐Led Evaluation of Generative AI Competence and Excellence Framework

Muhammet Güzelsoy, Anıl Erkan, Oğuzhan AKPINAR, Alper Keskin
article en

Abstract

OBJECTIVES: Interpretation of urodynamic studies is complex and subject to substantial inter-observer variability, even among experienced clinicians. While artificial intelligence has shown promise in urology, evidence regarding the performance of large language models (LLMs) in urodynamic trace interpretation remains limited. This study aimed to systematically compare the performance of multiple contemporary LLMs in interpreting standard urodynamic traces using an adapted ELEGANCE framework. METHODS: In this retrospective, single-centre comparative study, 119 anonymised urodynamic machine printouts retrieved from the institutional patient archive were independently interpreted by five LLM-based platforms (ChatGPT-5.2, Gemini 3, Perplexity AI Pro, DeepSeek-V3.2, and Grok 4.1). Case-specific expert reference interpretations were produced by one of three urologists, whereas all LLM-generated interpretations were independently evaluated by all three urologists. LLM-generated interpretations were assessed using an adapted version of the "Expert-Led Evaluation of Generative AI Competence and Excellence" (ELEGANCE) questionnaire, which evaluates relevance, completeness, applicability, structure, language/terminology, satisfaction, and hallucination. Between-model comparisons were performed using Friedman and post-hoc Wilcoxon signed-rank tests. RESULTS: [4] = 12.67, p = 0.013). Gemini achieved the highest mean score (25.77 ± 3.74), followed by Perplexity, DeepSeek, and ChatGPT-5.2, while Grok demonstrated significantly lower performance. Gemini consistently outperformed other models in clinically oriented domains, including relevance, completeness, and applicability (all p < 0.05). No hallucinations were identified in any model output. Inter-rater reliability among the three expert reviewers was good (average-measures ICC = 0.898, 95% CI 0.853-0.929; p < 0.001). CONCLUSION: Contemporary LLMs can generate structured interpretations of urodynamic traces, with meaningful performance differences across the evaluated platforms. While strengths in structure and terminology are evident, limitations in clinical reasoning persist. LLMs may serve as valuable assistive tools for standardisation and quality assurance in urodynamics but are not yet suitable for autonomous clinical decision-making.

Neurourology and Urodynamics
S.B.Ü. Bursa Yüksek İhtisas Eğitim ve Araştırma Hastanesi (TR)
Openalex Percentile: Top 18%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.