Entity-centric evaluation of large language model responses for medical question-answering tasks
Large language models (LLMs) are increasingly evaluated for clinical question-answering (QA) tasks, yet benchmark accuracy alone provides limited assurance of clinical reliability. Existing evaluation metrics often fail to capture whether model reasoning preserves patient-specific context and diagnostic intent, while many require external references or manual annotation, limiting scalability and real-world applicability. To address this gap, this study proposes E n t Q A , a reference-free, entity-centric metric that evaluates how well LLM-generated responses retain clinically relevant biomedical concepts from patient backgrounds and diagnostic questions. Across five medical QA benchmarks and seven Qwen 2.5 Instruct models (0.5B–72B parameters), E n t Q A showed consistently positive associations with model accuracy and scaling, outperforming conventional overlap- and embedding-based metrics that frequently exhibited weak or negative correlations. Group-level correlations with accuracy reached Spearman ρ = 0 . 9 2 8 6 , while correlations with model scale reached ρ = 0 . 2 5 2 . These findings suggest that E n t Q A provides a scalable and interpretable framework for assessing clinical fidelity and reasoning quality in healthcare LLMs without requiring gold standard references or external evidence.
Authors
- Vijaya B. Kolachalama (ORCID: https://orcid.org/0000-0002-5312-8644)
- Yi Liu
Institutions
- Boston University (US)
Publication Details
- Journal
- PLOS Digital Health
- Published
- 2026-10-06
- DOI
- https://doi.org/10.1371/journal.pdig.0001752
- Primary Topic
- Artificial Intelligence in Healthcare and Education
- Type
- article
- Field-Weighted Citation Impact
- 0.00