Automating Clavien–Dindo classification with large language models in percutaneous nephrolithotomy: an exploratory study

Abstract Purpose To explore the use of large language models for Clavien-Dindo classification of complications following percutaneous nephrolithotomy by benchmarking their performance against physician grading. Methods We evaluated ChatGPT, Copilot, DeepSeek, Gemini, Grok and Llama for defining Clavien-Dindo grades, classifying complications according to the Clavien-Dindo classification, and classifying case summaries with and without in-context prompting. Vignettes were obtained from a published consensus categorization of percutaneous nephrolithotomy complications. We measured agreement with Gwet’s AC2 coefficient and compared it with attending urologists and urology residents. Results All models accurately defined the Clavien-Dindo classification and graded with almost perfect agreement complication descriptions and case summaries (0.97 to 0.98 overall agreement). Six attendings and 15 residents classified the same vignettes, demonstrating lower agreement than the chatbots (0.95, p < 0.001 for residents; 0.96, p = 0.03 for attendings; 0.95, p < 0.001 for all participants). Models’ latency ranged from 0.77 to 3.90 min, while the fastest physician required 17 min to complete the same tasks. Models’ intrarater reliability was almost perfect: from 0.97 to 0.99. Interrater reliability between chatbots (0.97, 95% confidence interval: 0.97–0.98) was significantly higher than among physicians (0.93, 95% confidence interval: 0.92–0.94). Conclusion Large language models show promise in being faster, more accurate, and precise than physicians in Clavien-Dindo classification in an experimental setting. We support the need for a prospective agreement study using real-world clinical records. We urge discussion of the legal and ethical implications.

Authors

Institutions

Publication Details

Journal
World Journal of Urology
Published
2026-09-28
DOI
https://doi.org/10.1007/s00345-026-06833-z
Primary Topic
Kidney Stones and Urolithiasis Treatments
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Automating Clavien–Dindo classification with large language models in percutaneous nephrolithotomy: an exploratory study

William Carlos Nahas, Giovanni Scala Marchini, Eduardo Mazzucchi, Fábio César Miranda Torricelli et al.
World Journal of Urology
Kidney Stones and Urolithiasis Treatments
article

Automating Clavien–Dindo classification with large language models in percutaneous nephrolithotomy: an exploratory study

William Carlos Nahas, Giovanni Scala Marchini, Eduardo Mazzucchi, Fábio César Miranda Torricelli, Fabio Carvalho Vicentini, Alexandre Danilovic, Carlos Batagello, Rodrigo Perrella, Gustavo Perrone
article en

Abstract

Abstract Purpose To explore the use of large language models for Clavien-Dindo classification of complications following percutaneous nephrolithotomy by benchmarking their performance against physician grading. Methods We evaluated ChatGPT, Copilot, DeepSeek, Gemini, Grok and Llama for defining Clavien-Dindo grades, classifying complications according to the Clavien-Dindo classification, and classifying case summaries with and without in-context prompting. Vignettes were obtained from a published consensus categorization of percutaneous nephrolithotomy complications. We measured agreement with Gwet’s AC2 coefficient and compared it with attending urologists and urology residents. Results All models accurately defined the Clavien-Dindo classification and graded with almost perfect agreement complication descriptions and case summaries (0.97 to 0.98 overall agreement). Six attendings and 15 residents classified the same vignettes, demonstrating lower agreement than the chatbots (0.95, p < 0.001 for residents; 0.96, p = 0.03 for attendings; 0.95, p < 0.001 for all participants). Models’ latency ranged from 0.77 to 3.90 min, while the fastest physician required 17 min to complete the same tasks. Models’ intrarater reliability was almost perfect: from 0.97 to 0.99. Interrater reliability between chatbots (0.97, 95% confidence interval: 0.97–0.98) was significantly higher than among physicians (0.93, 95% confidence interval: 0.92–0.94). Conclusion Large language models show promise in being faster, more accurate, and precise than physicians in Clavien-Dindo classification in an experimental setting. We support the need for a prospective agreement study using real-world clinical records. We urge discussion of the legal and ethical implications.

World Journal of UrologyVol. 44(1)
Universidade de São Paulo (BR), Hospital das Clínicas da Faculdade de Medicina da Universidade de São Paulo (BR)
Quality Education
Openalex Percentile: Top 12%
Kidney Stones and Urolithiasis Treatments
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.