Evaluating Empathy in Large Language Model Responses Under Different Prompting Conditions: A Brief Literature Review

Large Language Models (LLMs) are increasingly used to produce conversational responses in emotionally sensitive settings. This paper presents a focused narrative review of research on empathetic communication in LLMs, with particular attention to prompting strategies and methods for evaluating empathy. Eight scholarly sources were reviewed, covering chatbot–human response comparisons, human–AI collaboration, mental-health applications, psychotherapy-informed prompting, persona-based prompting, empathy benchmarking, and the reliability of empathy evaluation. The reviewed evidence shows that LLMs can generate responses perceived as empathetic and that prompt design can influence emotional recognition, validation, and supportive language. The literature also indicates that persona instructions can alter model behavior in unintended ways and that empathy is difficult to represent with one universal metric. Multi-dimensional evaluation and comparison with human expert judgments therefore remain important when assessing empathetic communication. A further gap concerns consistency: existing research provides useful evidence about response quality and prompt effects, but fewer studies directly examine whether an LLM behaves consistently when the same scenario is repeated under the same prompting condition. The review concludes that prompting strategy and evaluation methodology should be studied together, particularly when LLMs are considered for emotionally sensitive communication.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-03
DOI
https://doi.org/10.5281/zenodo.23123174
Primary Topic
Digital Mental Health Interventions
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Evaluating Empathy in Large Language Model Responses Under Different Prompting Conditions: A Brief Literature Review

Dhanushraj Thangadurai Nadar
Zenodo (CERN European Organization for Nuclear Research)
Digital Mental Health Interventions
preprint

Evaluating Empathy in Large Language Model Responses Under Different Prompting Conditions: A Brief Literature Review

Dhanushraj Thangadurai Nadar
preprint en

Abstract

Large Language Models (LLMs) are increasingly used to produce conversational responses in emotionally sensitive settings. This paper presents a focused narrative review of research on empathetic communication in LLMs, with particular attention to prompting strategies and methods for evaluating empathy. Eight scholarly sources were reviewed, covering chatbot–human response comparisons, human–AI collaboration, mental-health applications, psychotherapy-informed prompting, persona-based prompting, empathy benchmarking, and the reliability of empathy evaluation. The reviewed evidence shows that LLMs can generate responses perceived as empathetic and that prompt design can influence emotional recognition, validation, and supportive language. The literature also indicates that persona instructions can alter model behavior in unintended ways and that empathy is difficult to represent with one universal metric. Multi-dimensional evaluation and comparison with human expert judgments therefore remain important when assessing empathetic communication. A further gap concerns consistency: existing research provides useful evidence about response quality and prompt effects, but fewer studies directly examine whether an LLM behaves consistently when the same scenario is repeated under the same prompting condition. The review concludes that prompting strategy and evaluation methodology should be studied together, particularly when LLMs are considered for emotionally sensitive communication.

Zenodo (CERN European Organization for Nuclear Research)
Digital Mental Health Interventions
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.