A Comparative Study of Generative Artificial Intelligence Versus Clinical Experts for Evidence-Based Decision-Making in Austere Environments

INTRODUCTION: The exponential growth of biomedical data hinders the rapid adoption of evidence-based practices in the Military Health System. This is critical in austere environments where teams rely on limited resources. Generative artificial intelligence (AI) offers a potential solution for rapidly synthesizing evidence. This comparative study evaluated the utility and feasibility of generative AI tools compared to human clinical experts in identifying protocols for surgical instrument reprocessing in austere settings. MATERIALS AND METHODS: We conducted a descriptive comparative study to query four AI platforms (NIPRGPT, ChatGPT, Google Gemini, GenAI.mil) and two clinical experts. The authors prompted each group to identify a single best recommendation for reprocessing surgical instruments in austere environments without steam sterilization capabilities. We compared the outputs based on time-to-completion, accessibility behind Department of War (DoW) firewalls, and clinical validity against a literature review. RESULTS: AI platforms generated recommendations in under 10 minutes. Clinical experts required 14 hours to review and synthesize data. Regarding accessibility, commercial platforms (ChatGPT, Gemini) were blocked by DoD firewalls, while GenAI.mil was accessible. Clinical experts recommended chlorine dioxide (ClO2) due to its sporicidal properties, which ensure sterility assurance. Only ChatGPT matched this recommendation. Conversely, GenAI.mil and Gemini recommended ortho-phthalaldehyde (OPA), and NIPRGPT recommended glutaraldehyde. The AI models prioritized processing speed over sterility assurance. CONCLUSIONS: Generative AI significantly reduces the cognitive load and time required to synthesize clinical protocols. However, government-hosted AI tools prioritized logistical factors over safety standards in this study. We identified an accessibility-accuracy paradox where the most accessible tool provided less rigorous safety recommendations. Implementation requires human verification and specific governance to ensure AI supports rather than replaces clinical judgment.

Authors

Institutions

Publication Details

Journal
Military Medicine
Published
2026-07-07
DOI
https://doi.org/10.1093/milmed/usag312
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

A Comparative Study of Generative Artificial Intelligence Versus Clinical Experts for Evidence-Based Decision-Making in Austere Environments

Christopher H. Stucky, Bethany I Atwood, Ross M Scallan, Chandler H. Moser et al.
Military Medicine
Artificial Intelligence in Healthcare and Education
article

A Comparative Study of Generative Artificial Intelligence Versus Clinical Experts for Evidence-Based Decision-Making in Austere Environments

Christopher H. Stucky, Bethany I Atwood, Ross M Scallan, Chandler H. Moser, Gina L Eberhardt, Stephanie Kessinger
article en

Abstract

INTRODUCTION: The exponential growth of biomedical data hinders the rapid adoption of evidence-based practices in the Military Health System. This is critical in austere environments where teams rely on limited resources. Generative artificial intelligence (AI) offers a potential solution for rapidly synthesizing evidence. This comparative study evaluated the utility and feasibility of generative AI tools compared to human clinical experts in identifying protocols for surgical instrument reprocessing in austere settings. MATERIALS AND METHODS: We conducted a descriptive comparative study to query four AI platforms (NIPRGPT, ChatGPT, Google Gemini, GenAI.mil) and two clinical experts. The authors prompted each group to identify a single best recommendation for reprocessing surgical instruments in austere environments without steam sterilization capabilities. We compared the outputs based on time-to-completion, accessibility behind Department of War (DoW) firewalls, and clinical validity against a literature review. RESULTS: AI platforms generated recommendations in under 10 minutes. Clinical experts required 14 hours to review and synthesize data. Regarding accessibility, commercial platforms (ChatGPT, Gemini) were blocked by DoD firewalls, while GenAI.mil was accessible. Clinical experts recommended chlorine dioxide (ClO2) due to its sporicidal properties, which ensure sterility assurance. Only ChatGPT matched this recommendation. Conversely, GenAI.mil and Gemini recommended ortho-phthalaldehyde (OPA), and NIPRGPT recommended glutaraldehyde. The AI models prioritized processing speed over sterility assurance. CONCLUSIONS: Generative AI significantly reduces the cognitive load and time required to synthesize clinical protocols. However, government-hosted AI tools prioritized logistical factors over safety standards in this study. We identified an accessibility-accuracy paradox where the most accessible tool provided less rigorous safety recommendations. Implementation requires human verification and specific governance to ensure AI supports rather than replaces clinical judgment.

Military Medicine
Landstuhl Regional Medical Center (DE), Madigan Army Medical Center (US), Unity Hospital (US)
Peace, Justice and strong institutions
Openalex Percentile: Top 11%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.