Development and Clinical Validation of an Automated LLM Judge for Evaluating Perioperative Patient Questions

Background: Large language models (LLMs) are increasingly used in patient-facing clinical applications, creating a need for scalable methods to evaluate the accuracy, safety, and appropriateness of their responses. Although physician review remains the reference standard, it is resource-intensive and difficult to scale. Objective: To develop and clinically validate an LLM-as-a-judge framework for evaluating responses to patients’ periprocedural questions. Methods: We retrospectively evaluated 1285 patient–AIVA interactions from 112 patients across two Mayo Clinic sites and multiple surgical specialties. Physician reviewers classified interactions using a predefined true-positive (TP), false-negative (FN), true-negative (TN), and false-positive (FP) framework. The finalized LLM judge independently evaluated the same interactions. Agreement was assessed using four-class and category-specific agreement, Cohen’s kappa, and a secondary binary analysis of response correctness. Results: Overall, four-class agreement was 93.1% (95% CI, 91.7–94.5%), with Cohen’s κ = 0.852 (95% CI, 0.823–0.882). Category-specific agreement was 94.1% for TP, 90.7% for FN, 95.5% for TN, and 84.9% for FP classifications. In the binary correctness analysis, accuracy was 93.4% (95% CI, 92.0–94.7%), sensitivity 94.6%, specificity 90.2%, precision 96.2%, and F1-score 0.954. McNemar’s test demonstrated no significant asymmetry between paired classifications (p = 0.11). Conclusions: A clinically grounded LLM judge closely approximated physician evaluation of patient-facing AI responses. These findings support its potential as a scalable assistive monitoring tool while preserving physician oversight for uncertain or safety-sensitive interactions.

Authors

Institutions

Publication Details

Journal
Bioengineering
Published
2026-09-24
DOI
https://doi.org/10.3390/bioengineering13101115
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Development and Clinical Validation of an Automated LLM Judge for Evaluating Perioperative Patient Questions

Bradley C. Leibovich, Ashton L. Boon, Bernardo Gabriele Collaço, Antonio J. Forte et al.
Bioengineering
Artificial Intelligence in Healthcare and Education
article

Development and Clinical Validation of an Automated LLM Judge for Evaluating Perioperative Patient Questions

Bradley C. Leibovich, Ashton L. Boon, Bernardo Gabriele Collaço, Antonio J. Forte, Carina Rosa Malena, Yunguo Yu, Nadia G. Wood, Prabha Srinivasagam, Zhihui Fang, Anjali Bhagra
article en

Abstract

Background: Large language models (LLMs) are increasingly used in patient-facing clinical applications, creating a need for scalable methods to evaluate the accuracy, safety, and appropriateness of their responses. Although physician review remains the reference standard, it is resource-intensive and difficult to scale. Objective: To develop and clinically validate an LLM-as-a-judge framework for evaluating responses to patients’ periprocedural questions. Methods: We retrospectively evaluated 1285 patient–AIVA interactions from 112 patients across two Mayo Clinic sites and multiple surgical specialties. Physician reviewers classified interactions using a predefined true-positive (TP), false-negative (FN), true-negative (TN), and false-positive (FP) framework. The finalized LLM judge independently evaluated the same interactions. Agreement was assessed using four-class and category-specific agreement, Cohen’s kappa, and a secondary binary analysis of response correctness. Results: Overall, four-class agreement was 93.1% (95% CI, 91.7–94.5%), with Cohen’s κ = 0.852 (95% CI, 0.823–0.882). Category-specific agreement was 94.1% for TP, 90.7% for FN, 95.5% for TN, and 84.9% for FP classifications. In the binary correctness analysis, accuracy was 93.4% (95% CI, 92.0–94.7%), sensitivity 94.6%, specificity 90.2%, precision 96.2%, and F1-score 0.954. McNemar’s test demonstrated no significant asymmetry between paired classifications (p = 0.11). Conclusions: A clinically grounded LLM judge closely approximated physician evaluation of patient-facing AI responses. These findings support its potential as a scalable assistive monitoring tool while preserving physician oversight for uncertain or safety-sensitive interactions.

BioengineeringVol. 13(10)
Mayo Clinic (US), Mayo Clinic in Arizona (US), Mayo Clinic in Florida (US)
Openalex Percentile: Top 15%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.