Development and Clinical Validation of an Automated LLM Judge for Evaluating Perioperative Patient Questions
Background: Large language models (LLMs) are increasingly used in patient-facing clinical applications, creating a need for scalable methods to evaluate the accuracy, safety, and appropriateness of their responses. Although physician review remains the reference standard, it is resource-intensive and difficult to scale. Objective: To develop and clinically validate an LLM-as-a-judge framework for evaluating responses to patients’ periprocedural questions. Methods: We retrospectively evaluated 1285 patient–AIVA interactions from 112 patients across two Mayo Clinic sites and multiple surgical specialties. Physician reviewers classified interactions using a predefined true-positive (TP), false-negative (FN), true-negative (TN), and false-positive (FP) framework. The finalized LLM judge independently evaluated the same interactions. Agreement was assessed using four-class and category-specific agreement, Cohen’s kappa, and a secondary binary analysis of response correctness. Results: Overall, four-class agreement was 93.1% (95% CI, 91.7–94.5%), with Cohen’s κ = 0.852 (95% CI, 0.823–0.882). Category-specific agreement was 94.1% for TP, 90.7% for FN, 95.5% for TN, and 84.9% for FP classifications. In the binary correctness analysis, accuracy was 93.4% (95% CI, 92.0–94.7%), sensitivity 94.6%, specificity 90.2%, precision 96.2%, and F1-score 0.954. McNemar’s test demonstrated no significant asymmetry between paired classifications (p = 0.11). Conclusions: A clinically grounded LLM judge closely approximated physician evaluation of patient-facing AI responses. These findings support its potential as a scalable assistive monitoring tool while preserving physician oversight for uncertain or safety-sensitive interactions.
Authors
- Bradley C. Leibovich (ORCID: https://orcid.org/0000-0003-0043-6659)
- Ashton L. Boon
- Bernardo Gabriele Collaço (ORCID: https://orcid.org/0000-0003-3845-2646)
- Antonio J. Forte (ORCID: https://orcid.org/0000-0003-2004-7538)
- Carina Rosa Malena (ORCID: https://orcid.org/0009-0003-2883-957X)
- Yunguo Yu (ORCID: https://orcid.org/0009-0006-2003-3813)
- Nadia G. Wood (ORCID: https://orcid.org/0000-0002-2685-6427)
- Prabha Srinivasagam
- Zhihui Fang (ORCID: https://orcid.org/0009-0000-4745-4745)
- Anjali Bhagra
Institutions
- Mayo Clinic (US)
- Mayo Clinic in Arizona (US)
- Mayo Clinic in Florida (US)
Publication Details
- Journal
- Bioengineering
- Published
- 2026-09-24
- DOI
- https://doi.org/10.3390/bioengineering13101115
- Primary Topic
- Artificial Intelligence in Healthcare and Education
- Type
- article
- Field-Weighted Citation Impact
- 0.00