Evaluating Empathy in Large Language Model Responses Under Different Prompting Conditions: A Brief Literature Review
Large Language Models (LLMs) are increasingly used to produce conversational responses in emotionally sensitive settings. This paper presents a focused narrative review of research on empathetic communication in LLMs, with particular attention to prompting strategies and methods for evaluating empathy. Eight scholarly sources were reviewed, covering chatbot–human response comparisons, human–AI collaboration, mental-health applications, psychotherapy-informed prompting, persona-based prompting, empathy benchmarking, and the reliability of empathy evaluation. The reviewed evidence shows that LLMs can generate responses perceived as empathetic and that prompt design can influence emotional recognition, validation, and supportive language. The literature also indicates that persona instructions can alter model behavior in unintended ways and that empathy is difficult to represent with one universal metric. Multi-dimensional evaluation and comparison with human expert judgments therefore remain important when assessing empathetic communication. A further gap concerns consistency: existing research provides useful evidence about response quality and prompt effects, but fewer studies directly examine whether an LLM behaves consistently when the same scenario is repeated under the same prompting condition. The review concludes that prompting strategy and evaluation methodology should be studied together, particularly when LLMs are considered for emotionally sensitive communication.
Authors
- Dhanushraj Thangadurai Nadar
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-03
- DOI
- https://doi.org/10.5281/zenodo.23123174
- Primary Topic
- Digital Mental Health Interventions
- Type
- preprint