Assessment of Clinical Empathy by an LLM-as-a-Judge: Bias Patterns in ATENA and a Replicable Audit Protocol for Medical Simulation

Background: Automated assessment of relational competencies by large language models (LLMs) is gaining traction in medical education, yet its empirical reliability remains poorly documented. Methods: We evaluated ATENA, a GPT-3.5-Turbo virtual-patient simulator deployed at the Faculty of Medicine of Nice (France), in two sequential components. First, 21 university hospital professors rated seven dimensions of realism on 7-point Likert scales. Second, across 129 student simulations of a single diagnostic disclosure scenario, expert human scores based on the empathic communication coding system (ECCS) were compared with ATENA’s automated scores using agreement metrics, Bland–Altman analysis, confusion matrix, and qualitative subcorpus analysis. Results: Realism was rated high on six of seven dimensions. Evaluative agreement was poor (ICC = 0.146), with a systematic overscoring bias (+0.80 points) and wide limits of agreement [−1.57; +3.16]. Three bias patterns emerged—score compression (76% of scores within 4.2–4.3), overscoring of weak performance, and underscoring of exemplary performance—indicating that ATENA responds to formal markers of empathy present in nearly every transcript. Conclusions: In this first-generation version of ATENA, the system distinguishes stronger from weaker interviews without being able to place them at the level assigned by expert human raters. These patterns preclude any certificatory use and support a critical, formative-only deployment grounded in instructor algorithmic literacy. The audit protocol applies to any LLM evaluator scored against human coding on a validated instrument.

Authors

Institutions

Publication Details

Journal
Education Sciences
Published
2026-09-14
DOI
https://doi.org/10.3390/educsci16091505
Primary Topic
Simulation-Based Education in Healthcare
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Assessment of Clinical Empathy by an LLM-as-a-Judge: Bias Patterns in ATENA and a Replicable Audit Protocol for Medical Simulation

Cyril Drouot, Alain Percivalle
Education Sciences
Simulation-Based Education in Healthcare
article

Assessment of Clinical Empathy by an LLM-as-a-Judge: Bias Patterns in ATENA and a Replicable Audit Protocol for Medical Simulation

Cyril Drouot, Alain Percivalle
article en

Abstract

Background: Automated assessment of relational competencies by large language models (LLMs) is gaining traction in medical education, yet its empirical reliability remains poorly documented. Methods: We evaluated ATENA, a GPT-3.5-Turbo virtual-patient simulator deployed at the Faculty of Medicine of Nice (France), in two sequential components. First, 21 university hospital professors rated seven dimensions of realism on 7-point Likert scales. Second, across 129 student simulations of a single diagnostic disclosure scenario, expert human scores based on the empathic communication coding system (ECCS) were compared with ATENA’s automated scores using agreement metrics, Bland–Altman analysis, confusion matrix, and qualitative subcorpus analysis. Results: Realism was rated high on six of seven dimensions. Evaluative agreement was poor (ICC = 0.146), with a systematic overscoring bias (+0.80 points) and wide limits of agreement [−1.57; +3.16]. Three bias patterns emerged—score compression (76% of scores within 4.2–4.3), overscoring of weak performance, and underscoring of exemplary performance—indicating that ATENA responds to formal markers of empathy present in nearly every transcript. Conclusions: In this first-generation version of ATENA, the system distinguishes stronger from weaker interviews without being able to place them at the level assigned by expert human raters. These patterns preclude any certificatory use and support a critical, formative-only deployment grounded in instructor algorithmic literacy. The audit protocol applies to any LLM evaluator scored against human coding on a validated instrument.

Education SciencesVol. 16(9)
Centre Hospitalier Universitaire de Nice (FR), Hôpital Pasteur (FR), Cognition Behaviour Technology (FR)
Quality Education
Openalex Percentile: Top 11%
Simulation-Based Education in Healthcare
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Assessment of Clinical Empathy by an LLM-as-a-Judge: Bias Patterns in ATENA and a Replicable Audit Protocol for Medical Simulation — Cyril Drouot, Alain Percivalle · Education Sciences (2026) | TGRS Research Map | TGRS