Evaluating reasoning-tuned large language models for clinical decision-making in spine surgery
Abstract Purpose Most clinical evaluations of large language models assess factual recall rather than the multi-step reasoning behind operative plans. Reasoning-tuned models, post-trained to generate explicit intermediate reasoning before answering, may better approximate surgical decision-making. We compared two such models from different development ecosystems, OpenAI o1 and DeepSeek R1, on complex spine cases. Methods Ten synthetic vignettes spanning deformity, degenerative, urgent, and cervical pathology were analysed; an eleventh was excluded owing to a survey error. Both models received identical prompts. Outputs were anonymised and randomised. Eight fellowship-trained spine surgeons (five consultants, three fellows) scored diagnostic accuracy, reasoning and thoroughness, surgical plan appropriateness, and clarity on 5-point Likert scales, giving 78 paired evaluations. Analysis used linear mixed-effects models with crossed random intercepts for rater and vignette (Bonferroni-adjusted alpha 0.0125); reliability was assessed by intraclass correlation. Results o1 scored higher in all four domains, significantly for diagnostic accuracy (4.6 vs. 4.3; mean difference 0.24, 95% CI 0.11–0.38) and clarity (4.4 vs. 4.2; 0.26, 0.07–0.44); reasoning and thoroughness (0.21, 0.03–0.38), and surgical plan appropriateness (0.17, − 0.03 to 0.36) did not meet the adjusted threshold. Inter-rater reliability was poor (ICC 0.03–0.07). Clarity correlated strongly with the clinical domains (rho 0.74–0.79); adjusting for clarity attenuated the o1 advantage by 55% to over 100%, leaving no detectable clinical-domain difference. Conclusion Both models produced plans surgeons usually rated good or excellent. o1 was rated higher, significantly only for diagnostic accuracy and clarity, most of it reflecting presentation rather than clinical content. Surgeon oversight and format-normalised evaluation designs remain essential.
Authors
- David L. Skaggs (ORCID: https://orcid.org/0000-0001-6137-5576)
- CW Lam (ORCID: https://orcid.org/0009-0000-1970-3038)
- Morgan Jones (ORCID: https://orcid.org/0000-0002-6953-1091)
- Adit Ravishankar
- George McKay
- Alex Bulloso
- Conor T. Boylan
Institutions
- Cedars-Sinai Medical Center (US)
- University of Liverpool (GB)
- University of Cambridge (GB)
- Royal Orthopaedic Hospital (GB)
- Alder Hey Children's NHS Foundation Trust (GB)
- Alder Hey Children's Hospital (GB)
Publication Details
- Journal
- Spine Deformity
- Published
- 2026-09-28
- DOI
- https://doi.org/10.1007/s43390-026-01557-x
- Primary Topic
- Clinical Reasoning and Diagnostic Skills
- Type
- article
- Field-Weighted Citation Impact
- 0.00