From structured analysis to normative decision-making: a comparative evaluation of large language models in clinical ethics consultation

Although large language models are increasingly used in healthcare, evidence on their role in clinical ethics consultation remains limited. This study assessed their ability to generate consultation reports, their sensitivity to context and normative pluralism, and the conditions under which their use may be ethically defensible. Four published clinical cases were presented to ChatGPT (GPT-4-based), Gemini (1.5 series), and DeepSeek (V2-based) using a standardized Turkish prompt. One output per model per case produced 12 reports. Four medical ethics experts, blinded to model identity, evaluated the reports using five criteria, and the scores were statistically compared. A post hoc exploratory qualitative analysis examined the accuracy, verifiability, currency, and case relevance of legal and ethical references, and was extended to individual and local contextual elements and normative orientations. All models identified clinical and social dimensions, recognized ethical problems, and produced structured reports. DeepSeek scored 392/400, ChatGPT 383/400, and Gemini 381/400, with no statistically significant differences. Reports included patient preferences, decision-making capacity, surrogate decision-making, and caregiver concerns, but did not always show how these elements were weighed against conflicting values. Similar normative orientations suggested a possible narrowing of normative diversity. Only DeepSeek extensively cited national legal and ethical frameworks; most references were inaccurate, contextually inappropriate, or unverifiable. These problems were missed during expert evaluation and the authors’ initial review, and the references were initially perceived as a strength. Large language models may support the formal and analytical components of clinical ethics consultation, but high scores do not establish context-sensitive, pluralistic, or reliable normative reasoning. Convergence may reduce the visibility of ethically reasonable alternatives. The failure to detect inaccurate references despite human oversight shows the limits of individual review. Outputs should be treated as unverified preliminary assessments, with legal and ethical claims checked against primary sources and use supported by institutional oversight, local context sensitivity, and safeguards for normative pluralism.

Authors

Institutions

Publication Details

Journal
BMC Medical Ethics
Published
2026-10-07
DOI
https://doi.org/10.1186/s12910-026-01626-w
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

From structured analysis to normative decision-making: a comparative evaluation of large language models in clinical ethics consultation

Özge Pasin, İsmail Topçu, Beyzanur KAÇ, Merve ERDEM
BMC Medical Ethics
Artificial Intelligence in Healthcare and Education
article

From structured analysis to normative decision-making: a comparative evaluation of large language models in clinical ethics consultation

Özge Pasin, İsmail Topçu, Beyzanur KAÇ, Merve ERDEM
article en

Abstract

Although large language models are increasingly used in healthcare, evidence on their role in clinical ethics consultation remains limited. This study assessed their ability to generate consultation reports, their sensitivity to context and normative pluralism, and the conditions under which their use may be ethically defensible. Four published clinical cases were presented to ChatGPT (GPT-4-based), Gemini (1.5 series), and DeepSeek (V2-based) using a standardized Turkish prompt. One output per model per case produced 12 reports. Four medical ethics experts, blinded to model identity, evaluated the reports using five criteria, and the scores were statistically compared. A post hoc exploratory qualitative analysis examined the accuracy, verifiability, currency, and case relevance of legal and ethical references, and was extended to individual and local contextual elements and normative orientations. All models identified clinical and social dimensions, recognized ethical problems, and produced structured reports. DeepSeek scored 392/400, ChatGPT 383/400, and Gemini 381/400, with no statistically significant differences. Reports included patient preferences, decision-making capacity, surrogate decision-making, and caregiver concerns, but did not always show how these elements were weighed against conflicting values. Similar normative orientations suggested a possible narrowing of normative diversity. Only DeepSeek extensively cited national legal and ethical frameworks; most references were inaccurate, contextually inappropriate, or unverifiable. These problems were missed during expert evaluation and the authors’ initial review, and the references were initially perceived as a strength. Large language models may support the formal and analytical components of clinical ethics consultation, but high scores do not establish context-sensitive, pluralistic, or reliable normative reasoning. Convergence may reduce the visibility of ethically reasonable alternatives. The failure to detect inaccurate references despite human oversight shows the limits of individual review. Outputs should be treated as unverified preliminary assessments, with legal and ethical claims checked against primary sources and use supported by institutional oversight, local context sensitivity, and safeguards for normative pluralism.

BMC Medical Ethics
University of Health Science (KH), Sağlık Bilimleri Üniversitesi (TR), Maltepe University (TR), University of Health Sciences Antigua (AG)
Openalex Percentile: Top 19%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.