Automating the triage of rheumatology outpatient referrals: a comparative evaluation of 23 large language models under simple and advanced prompting

Abstract Triage of rheumatology outpatient referrals is a high-volume administrative task that consumes senior specialist time without advancing patient care. The human triage system is only moderately accurate and reproducible. We assessed whether contemporary large language models (LLMs) are able to perform well enough to support automating this task in practice. In addition, the effects of different prompting techniques on triage accuracy and cost was assessed to help identify the optimal approach. Twenty referral scenarios spanning the urgency spectrum were created by a certified Australian rheumatologist. Four rheumatologists triaged all cases independently and blinded to produce a consensus reference standard; inter-rater reliability was quantified. Twenty-three LLMs each triaged every referral into one of five urgency categories, three times (1380 outputs per referral). The experiment was run with a simple prompt and repeated with an advanced prompt supplying explicit triage expectations and worked examples. In line with published literature, agreement among the four rheumatologists was moderate (Fleiss kappa 0.60; mean pairwise linear-weighted kappa 0.79), with all four assigning categories within one level of each other in 90% of cases. All 2760 model calls returned an interpretable result. Under the simple prompt, performance separated into distinct tiers, larger models were more accurate (Spearman rho = 0.42; P =.047) and accuracy tracked cost. Advanced prompting reduced between-model variance in accuracy 5.3-fold (0.014 to 0.003; Levene P =.01), abolished the size-accuracy association (rho=-0.05; P =.83) and removed the accuracy-cost relationship. The best models matched expert consensus on 18 of 20 cases, comparable to the reproducibility observed among the rheumatologists themselves. Under-triage errors persisted with some LLMs. Contemporary LLMs categorised rheumatology referral urgency with an accuracy and reproducibility comparable to that reported for human triage systems, although human triage was not tested in this study. Advanced prompting appeared to substitute for the reasoning capability of larger models, suggesting that adequate performance on this task may not require the most expensive models. These findings indicate that automation of this administrative task is technically feasible. Candidate models that warrant prospective evaluation are identified.

Authors

Institutions

Publication Details

Journal
Rheumatology International
Published
2026-10-09
DOI
https://doi.org/10.1007/s00296-026-06317-8
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Automating the triage of rheumatology outpatient referrals: a comparative evaluation of 23 large language models under simple and advanced prompting

Lynden J. Roberts
Rheumatology International
Artificial Intelligence in Healthcare and Education
article

Automating the triage of rheumatology outpatient referrals: a comparative evaluation of 23 large language models under simple and advanced prompting

Lynden J. Roberts
article en

Abstract

Abstract Triage of rheumatology outpatient referrals is a high-volume administrative task that consumes senior specialist time without advancing patient care. The human triage system is only moderately accurate and reproducible. We assessed whether contemporary large language models (LLMs) are able to perform well enough to support automating this task in practice. In addition, the effects of different prompting techniques on triage accuracy and cost was assessed to help identify the optimal approach. Twenty referral scenarios spanning the urgency spectrum were created by a certified Australian rheumatologist. Four rheumatologists triaged all cases independently and blinded to produce a consensus reference standard; inter-rater reliability was quantified. Twenty-three LLMs each triaged every referral into one of five urgency categories, three times (1380 outputs per referral). The experiment was run with a simple prompt and repeated with an advanced prompt supplying explicit triage expectations and worked examples. In line with published literature, agreement among the four rheumatologists was moderate (Fleiss kappa 0.60; mean pairwise linear-weighted kappa 0.79), with all four assigning categories within one level of each other in 90% of cases. All 2760 model calls returned an interpretable result. Under the simple prompt, performance separated into distinct tiers, larger models were more accurate (Spearman rho = 0.42; P =.047) and accuracy tracked cost. Advanced prompting reduced between-model variance in accuracy 5.3-fold (0.014 to 0.003; Levene P =.01), abolished the size-accuracy association (rho=-0.05; P =.83) and removed the accuracy-cost relationship. The best models matched expert consensus on 18 of 20 cases, comparable to the reproducibility observed among the rheumatologists themselves. Under-triage errors persisted with some LLMs. Contemporary LLMs categorised rheumatology referral urgency with an accuracy and reproducibility comparable to that reported for human triage systems, although human triage was not tested in this study. Advanced prompting appeared to substitute for the reasoning capability of larger models, suggesting that adequate performance on this task may not require the most expensive models. These findings indicate that automation of this administrative task is technically feasible. Candidate models that warrant prospective evaluation are identified.

Rheumatology InternationalVol. 46(10)
Monash Health (AU), Monash University (AU)
Openalex Percentile: Top 20%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.