Entity-centric evaluation of large language model responses for medical question-answering tasks

Large language models (LLMs) are increasingly evaluated for clinical question-answering (QA) tasks, yet benchmark accuracy alone provides limited assurance of clinical reliability. Existing evaluation metrics often fail to capture whether model reasoning preserves patient-specific context and diagnostic intent, while many require external references or manual annotation, limiting scalability and real-world applicability. To address this gap, this study proposes E n t Q A , a reference-free, entity-centric metric that evaluates how well LLM-generated responses retain clinically relevant biomedical concepts from patient backgrounds and diagnostic questions. Across five medical QA benchmarks and seven Qwen 2.5 Instruct models (0.5B–72B parameters), E n t Q A showed consistently positive associations with model accuracy and scaling, outperforming conventional overlap- and embedding-based metrics that frequently exhibited weak or negative correlations. Group-level correlations with accuracy reached Spearman ρ = 0 . 9 2 8 6 , while correlations with model scale reached ρ = 0 . 2 5 2 . These findings suggest that E n t Q A provides a scalable and interpretable framework for assessing clinical fidelity and reasoning quality in healthcare LLMs without requiring gold standard references or external evidence.

Authors

Institutions

Publication Details

Journal
PLOS Digital Health
Published
2026-10-06
DOI
https://doi.org/10.1371/journal.pdig.0001752
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Entity-centric evaluation of large language model responses for medical question-answering tasks

Vijaya B. Kolachalama, Yi Liu
PLOS Digital Health
Artificial Intelligence in Healthcare and Education
article

Entity-centric evaluation of large language model responses for medical question-answering tasks

Vijaya B. Kolachalama, Yi Liu
article en

Abstract

Large language models (LLMs) are increasingly evaluated for clinical question-answering (QA) tasks, yet benchmark accuracy alone provides limited assurance of clinical reliability. Existing evaluation metrics often fail to capture whether model reasoning preserves patient-specific context and diagnostic intent, while many require external references or manual annotation, limiting scalability and real-world applicability. To address this gap, this study proposes E n t Q A , a reference-free, entity-centric metric that evaluates how well LLM-generated responses retain clinically relevant biomedical concepts from patient backgrounds and diagnostic questions. Across five medical QA benchmarks and seven Qwen 2.5 Instruct models (0.5B–72B parameters), E n t Q A showed consistently positive associations with model accuracy and scaling, outperforming conventional overlap- and embedding-based metrics that frequently exhibited weak or negative correlations. Group-level correlations with accuracy reached Spearman ρ = 0 . 9 2 8 6 , while correlations with model scale reached ρ = 0 . 2 5 2 . These findings suggest that E n t Q A provides a scalable and interpretable framework for assessing clinical fidelity and reasoning quality in healthcare LLMs without requiring gold standard references or external evidence.

PLOS Digital HealthVol. 5(10)
Boston University (US)
Openalex Percentile: Top 19%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Entity-centric evaluation of large language model responses for medical question-answering tasks — Vijaya B. Kolachalama, Yi Liu · PLOS Digital Health (2026) | TGRS Research Map | TGRS