From structured analysis to normative decision-making: a comparative evaluation of large language models in clinical ethics consultation
Although large language models are increasingly used in healthcare, evidence on their role in clinical ethics consultation remains limited. This study assessed their ability to generate consultation reports, their sensitivity to context and normative pluralism, and the conditions under which their use may be ethically defensible. Four published clinical cases were presented to ChatGPT (GPT-4-based), Gemini (1.5 series), and DeepSeek (V2-based) using a standardized Turkish prompt. One output per model per case produced 12 reports. Four medical ethics experts, blinded to model identity, evaluated the reports using five criteria, and the scores were statistically compared. A post hoc exploratory qualitative analysis examined the accuracy, verifiability, currency, and case relevance of legal and ethical references, and was extended to individual and local contextual elements and normative orientations. All models identified clinical and social dimensions, recognized ethical problems, and produced structured reports. DeepSeek scored 392/400, ChatGPT 383/400, and Gemini 381/400, with no statistically significant differences. Reports included patient preferences, decision-making capacity, surrogate decision-making, and caregiver concerns, but did not always show how these elements were weighed against conflicting values. Similar normative orientations suggested a possible narrowing of normative diversity. Only DeepSeek extensively cited national legal and ethical frameworks; most references were inaccurate, contextually inappropriate, or unverifiable. These problems were missed during expert evaluation and the authors’ initial review, and the references were initially perceived as a strength. Large language models may support the formal and analytical components of clinical ethics consultation, but high scores do not establish context-sensitive, pluralistic, or reliable normative reasoning. Convergence may reduce the visibility of ethically reasonable alternatives. The failure to detect inaccurate references despite human oversight shows the limits of individual review. Outputs should be treated as unverified preliminary assessments, with legal and ethical claims checked against primary sources and use supported by institutional oversight, local context sensitivity, and safeguards for normative pluralism.
Authors
- Özge Pasin (ORCID: https://orcid.org/0000-0001-6530-0942)
- İsmail Topçu (ORCID: https://orcid.org/0000-0002-9572-1251)
- Beyzanur KAÇ (ORCID: https://orcid.org/0000-0002-7461-1942)
- Merve ERDEM (ORCID: https://orcid.org/0000-0002-8616-991X)
Institutions
- University of Health Science (KH)
- Sağlık Bilimleri Üniversitesi (TR)
- Maltepe University (TR)
- University of Health Sciences Antigua (AG)
Publication Details
- Journal
- BMC Medical Ethics
- Published
- 2026-10-07
- DOI
- https://doi.org/10.1186/s12910-026-01626-w
- Primary Topic
- Artificial Intelligence in Healthcare and Education
- Type
- article
- Field-Weighted Citation Impact
- 0.00