Evaluating reasoning-tuned large language models for clinical decision-making in spine surgery

Abstract Purpose Most clinical evaluations of large language models assess factual recall rather than the multi-step reasoning behind operative plans. Reasoning-tuned models, post-trained to generate explicit intermediate reasoning before answering, may better approximate surgical decision-making. We compared two such models from different development ecosystems, OpenAI o1 and DeepSeek R1, on complex spine cases. Methods Ten synthetic vignettes spanning deformity, degenerative, urgent, and cervical pathology were analysed; an eleventh was excluded owing to a survey error. Both models received identical prompts. Outputs were anonymised and randomised. Eight fellowship-trained spine surgeons (five consultants, three fellows) scored diagnostic accuracy, reasoning and thoroughness, surgical plan appropriateness, and clarity on 5-point Likert scales, giving 78 paired evaluations. Analysis used linear mixed-effects models with crossed random intercepts for rater and vignette (Bonferroni-adjusted alpha 0.0125); reliability was assessed by intraclass correlation. Results o1 scored higher in all four domains, significantly for diagnostic accuracy (4.6 vs. 4.3; mean difference 0.24, 95% CI 0.11–0.38) and clarity (4.4 vs. 4.2; 0.26, 0.07–0.44); reasoning and thoroughness (0.21, 0.03–0.38), and surgical plan appropriateness (0.17, − 0.03 to 0.36) did not meet the adjusted threshold. Inter-rater reliability was poor (ICC 0.03–0.07). Clarity correlated strongly with the clinical domains (rho 0.74–0.79); adjusting for clarity attenuated the o1 advantage by 55% to over 100%, leaving no detectable clinical-domain difference. Conclusion Both models produced plans surgeons usually rated good or excellent. o1 was rated higher, significantly only for diagnostic accuracy and clarity, most of it reflecting presentation rather than clinical content. Surgeon oversight and format-normalised evaluation designs remain essential.

Authors

Institutions

Publication Details

Journal
Spine Deformity
Published
2026-09-28
DOI
https://doi.org/10.1007/s43390-026-01557-x
Primary Topic
Clinical Reasoning and Diagnostic Skills
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Evaluating reasoning-tuned large language models for clinical decision-making in spine surgery

David L. Skaggs, CW Lam, Morgan Jones, Adit Ravishankar et al.
Spine Deformity
Clinical Reasoning and Diagnostic Skills
article

Evaluating reasoning-tuned large language models for clinical decision-making in spine surgery

David L. Skaggs, CW Lam, Morgan Jones, Adit Ravishankar, George McKay, Alex Bulloso, Conor T. Boylan
article en

Abstract

Abstract Purpose Most clinical evaluations of large language models assess factual recall rather than the multi-step reasoning behind operative plans. Reasoning-tuned models, post-trained to generate explicit intermediate reasoning before answering, may better approximate surgical decision-making. We compared two such models from different development ecosystems, OpenAI o1 and DeepSeek R1, on complex spine cases. Methods Ten synthetic vignettes spanning deformity, degenerative, urgent, and cervical pathology were analysed; an eleventh was excluded owing to a survey error. Both models received identical prompts. Outputs were anonymised and randomised. Eight fellowship-trained spine surgeons (five consultants, three fellows) scored diagnostic accuracy, reasoning and thoroughness, surgical plan appropriateness, and clarity on 5-point Likert scales, giving 78 paired evaluations. Analysis used linear mixed-effects models with crossed random intercepts for rater and vignette (Bonferroni-adjusted alpha 0.0125); reliability was assessed by intraclass correlation. Results o1 scored higher in all four domains, significantly for diagnostic accuracy (4.6 vs. 4.3; mean difference 0.24, 95% CI 0.11–0.38) and clarity (4.4 vs. 4.2; 0.26, 0.07–0.44); reasoning and thoroughness (0.21, 0.03–0.38), and surgical plan appropriateness (0.17, − 0.03 to 0.36) did not meet the adjusted threshold. Inter-rater reliability was poor (ICC 0.03–0.07). Clarity correlated strongly with the clinical domains (rho 0.74–0.79); adjusting for clarity attenuated the o1 advantage by 55% to over 100%, leaving no detectable clinical-domain difference. Conclusion Both models produced plans surgeons usually rated good or excellent. o1 was rated higher, significantly only for diagnostic accuracy and clarity, most of it reflecting presentation rather than clinical content. Surgeon oversight and format-normalised evaluation designs remain essential.

Spine Deformity
Cedars-Sinai Medical Center (US), University of Liverpool (GB), University of Cambridge (GB), Royal Orthopaedic Hospital (GB), Alder Hey Children's NHS Foundation Trust (GB), Alder Hey Children's Hospital (GB)
Peace, Justice and strong institutions
Openalex Percentile: Top 9%
Clinical Reasoning and Diagnostic Skills
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.