The economics of accuracy for medical reasoning with large language models

{"Deploying":[0],"large":[1],"language":[2],"models":[3,27],"(LLMs)":[4],"in":[5,266],"clinical":[6],"settings":[7],"is":[8,280],"limited":[9],"by":[10,290],"security,":[11],"reliability,":[12],"latency,":[13],"and":[14,36,52,74,96,101,113,116,126,151,176,189,229,251,316,323,328],"accessibility":[15],"concerns":[16],"that":[17,166],"favor":[18],"smaller,":[19],"on-device":[20],"or":[21,223],"on-premise":[22],"models.":[23],"However,":[24,68],"these":[25,44,72,118,295],"smaller":[26,199],"may":[28],"struggle":[29],"to":[30,65,180,240,261,305,325,335],"meet":[31],"accuracy":[32,175,201,258,289,327],"requirements.":[33],"While":[34],"fine-tuning":[35,236],"retrieval-augmented":[37],"generation":[38],"(RAG)":[39],"can":[40],"improve":[41],"domain-specific":[42],"accuracy,":[43],"methods":[45],"require":[46],"additional":[47],"labeled":[48],"data,":[49],"technical":[50],"skill,":[51],"infrastructure.":[53],"In":[54],"contrast,":[55],"test-time":[56],"scaling-allocating":[57],"extra":[58],"token-budget":[59],"during":[60],"inference-offers":[61],"a":[62,127,155,198,209,300],"training-free":[63],"alternative":[64],"increasing":[66],"accuracy.":[67],"the":[69,99,136,182,232,267],"trade-offs":[70,184,308],"between":[71],"strategies":[73],"their":[75],"interaction":[76],"with":[77,135,249,303],"model":[78,190],"size":[79],"remain":[80],"poorly":[81],"understood":[82],"for":[83,142,162,313,331],"medical":[84,132,235,314],"reasoning.":[85],"To":[86],"address":[87],"this":[88],"gap,":[89],"we":[90,159,298],"compare":[91],"three":[92],"approaches-test-time":[93],"scaling,":[94],"fine-tuning,":[95,188],"context":[97,279],"grounding-using":[98],"Gemma":[100],"MedGemma":[102,114,256],"family":[103],"of":[104,129,138,208,234,269],"LLMs":[105],"(Gemma-3":[106],"1B,":[107],"Gemma-3":[108,110],"4B,":[109],"27B,":[111],"MedGemma-4B,":[112],"27B)":[115],"evaluate":[117],"systems":[119,312],"across":[120,185],"common":[121],"biomedical":[122],"question-answering":[123],"(QA)":[124],"datasets":[125],"set":[128],"recently":[130],"released":[131],"exam":[133],"questions":[134],"performance":[137],"practicing":[139],"clinicians":[140],"available":[141],"comparison.":[143],"We":[144,173,192,271,318],"test":[145],"baseline":[146],"prompts":[147],"(direct":[148],"answer,":[149],"Chain-of-Thought,":[150],"self-consistency)":[152],"while":[153],"introducing":[154],"new":[156],"prompting":[157],"method":[158],"call":[160],"\\"prompt-chaining":[161],"continuous":[163],"reflection\\"":[164],"(PCCR)":[165],"forces":[167],"inference":[168],"time":[169],"minimum":[170],"token-generation":[171],"budgets.":[172],"assess":[174],"tokens-generated,":[177],"allowing":[178],"us":[179],"investigate":[181],"accuracy-efficiency":[183],"prompting,":[186],"context-grounding,":[187,222],"scales.":[191],"discover":[193],"equivalency":[194],"point":[195],"configurations":[196],"where":[197],"model's":[200,211],"falls":[202],"within":[203,213],"one":[204],"95%":[205,245],"confidence":[206],"interval":[207],"larger":[210],"(typically":[212],"1-4":[214],"percentage":[215,242,292],"points)":[216],"reached":[217],"through":[218],"increased":[219],"reasoning":[220,254,283,315],"budgets,":[221],"fine-tuning.":[224],"Specific":[225],"effects":[226],"are":[227],"apparent":[228],"statistically":[230],"supported:":[231],"benefit":[233],"grew":[237],"from":[238,259],"+4.6":[239],"+15.7":[241],"points":[243],"(non-overlapping":[244],"CIs)":[246],"when":[247,277,309],"paired":[248],"self-consistency,":[250],"enforced":[252],"extended":[253,282],"raised":[255],"27B":[257],"58.1%":[260],"80.1%":[262],"(p":[263],"<":[264],"10-5)":[265],"absence":[268],"context.":[270],"also":[272],"identify":[273],"an":[274],"\\"overthinking\\"":[275],"inflection:":[276],"high-quality":[278],"available,":[281],"beyond":[284],"roughly":[285],"128-256":[286],"tokens":[287],"degrades":[288],"7-13":[291],"points.":[293],"Using":[294],"empirical":[296],"results,":[297],"formulate":[299],"general":[301],"framework":[302],"equations":[304],"balance":[306],"cost-benefit":[307],"engineering":[310],"LLM-based":[311],"QA.":[317],"recommend":[319],"generalizable":[320],"configurations,":[321],"designs,":[322],"patterns":[324],"achieve":[326],"efficiency":[329],"objectives":[330],"example":[332],"use-cases":[333],"relevant":[334],"healthcare":[336],"organizations.":[337]}

Authors

Institutions

Publication Details

Journal
PLOS Digital Health
Published
2026-09-18
DOI
https://doi.org/10.1371/journal.pdig.0001182
Primary Topic
Topic Modeling
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

The economics of accuracy for medical reasoning with large language models

Kiran Bhattacharyya, Sreeram Kamabattula
PLOS Digital Health
Topic Modeling
article

The economics of accuracy for medical reasoning with large language models

Kiran Bhattacharyya, Sreeram Kamabattula
article en

Abstract

Deploying large language models (LLMs) in clinical settings is limited by security, reliability, latency, and accessibility concerns that favor smaller, on-device or on-premise models. However, these smaller models may struggle to meet accuracy requirements. While fine-tuning and retrieval-augmented generation (RAG) can improve domain-specific accuracy, these methods require additional labeled data, technical skill, and infrastructure. In contrast, test-time scaling-allocating extra token-budget during inference-offers a training-free alternative to increasing accuracy. However, the trade-offs between these strategies and their interaction with model size remain poorly understood for medical reasoning. To address this gap, we compare three approaches-test-time scaling, fine-tuning, and context grounding-using the Gemma and MedGemma family of LLMs (Gemma-3 1B, Gemma-3 4B, Gemma-3 27B, MedGemma-4B, and MedGemma 27B) and evaluate these systems across common biomedical question-answering (QA) datasets and a set of recently released medical exam questions with the performance of practicing clinicians available for comparison. We test baseline prompts (direct answer, Chain-of-Thought, and self-consistency) while introducing a new prompting method we call "prompt-chaining for continuous reflection" (PCCR) that forces inference time minimum token-generation budgets. We assess accuracy and tokens-generated, allowing us to investigate the accuracy-efficiency trade-offs across prompting, context-grounding, fine-tuning, and model scales. We discover equivalency point configurations where a smaller model's accuracy falls within one 95% confidence interval of a larger model's (typically within 1-4 percentage points) reached through increased reasoning budgets, context-grounding, or fine-tuning. Specific effects are apparent and statistically supported: the benefit of medical fine-tuning grew from +4.6 to +15.7 percentage points (non-overlapping 95% CIs) when paired with self-consistency, and enforced extended reasoning raised MedGemma 27B accuracy from 58.1% to 80.1% (p < 10-5) in the absence of context. We also identify an "overthinking" inflection: when high-quality context is available, extended reasoning beyond roughly 128-256 tokens degrades accuracy by 7-13 percentage points. Using these empirical results, we formulate a general framework with equations to balance cost-benefit trade-offs when engineering LLM-based systems for medical reasoning and QA. We recommend generalizable configurations, designs, and patterns to achieve accuracy and efficiency objectives for example use-cases relevant to healthcare organizations.

PLOS Digital HealthVol. 5(9)
Intuitive Surgical (Switzerland) (CH), SKA Telescope, South Africa (ZA), Intuitive Surgical (United States) (US)
Industry, innovation and infrastructure
Openalex Percentile: Top 9%
Topic Modeling
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.