Simulated Morality, Misplaced Trust: The Risks of Treating AI as a Moral Partner

Abstract The article argues that many contemporary AI alignment practices risk a mistaken assimilation of moral agency to statistical learning. Techniques such as reinforcement learning from human feedback and constitutional AI often treat morality as a behavioral function that can be approximated from human discourse, behavior, and large-scale interaction data. Drawing on Tomasello’s account of shared intentionality, Hegel’s theory of recognition, and the second-personal tradition in moral philosophy, the article contends that moral agency is not exhausted by norm-conforming output, but presupposes participation in a space of mutual accountability, justificatory practice, and normative self-binding. On this view, and given the architecture of present generative systems, the production of normatively fluent behavior and reason-like discourse is not sufficient for moral partnership: absent interpersonal identity, recognition, and second-personal answerability, such systems can at best simulate the outer form of moral conduct. The central risk, therefore, is not that these systems behave “immorally,” but that their reliable norm-conforming behavior is misread as evidence of genuine moral commitment, encouraging misplaced trust and over-ascription of responsibility. The article proposes a reframing of alignment from an ethical-pedagogical project to an institutional one: instead of presuming or attempting to cultivate artificial moral agents on the basis of behavioral proxies, governance should focus on designing legal, technical, and organizational structures that constrain and render machine behavior auditable and accountable, while keeping moral responsibility firmly anchored in human agents and institutions.

Authors

Publication Details

Journal
Philosophy & Technology
Published
2026-09-28
DOI
https://doi.org/10.1007/s13347-026-01180-8
Primary Topic
Ethics and Social Impacts of AI
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Simulated Morality, Misplaced Trust: The Risks of Treating AI as a Moral Partner

Saša Josifović
Philosophy & Technology
Ethics and Social Impacts of AI
article

Simulated Morality, Misplaced Trust: The Risks of Treating AI as a Moral Partner

Saša Josifović
article en

Abstract

Abstract The article argues that many contemporary AI alignment practices risk a mistaken assimilation of moral agency to statistical learning. Techniques such as reinforcement learning from human feedback and constitutional AI often treat morality as a behavioral function that can be approximated from human discourse, behavior, and large-scale interaction data. Drawing on Tomasello’s account of shared intentionality, Hegel’s theory of recognition, and the second-personal tradition in moral philosophy, the article contends that moral agency is not exhausted by norm-conforming output, but presupposes participation in a space of mutual accountability, justificatory practice, and normative self-binding. On this view, and given the architecture of present generative systems, the production of normatively fluent behavior and reason-like discourse is not sufficient for moral partnership: absent interpersonal identity, recognition, and second-personal answerability, such systems can at best simulate the outer form of moral conduct. The central risk, therefore, is not that these systems behave “immorally,” but that their reliable norm-conforming behavior is misread as evidence of genuine moral commitment, encouraging misplaced trust and over-ascription of responsibility. The article proposes a reframing of alignment from an ethical-pedagogical project to an institutional one: instead of presuming or attempting to cultivate artificial moral agents on the basis of behavioral proxies, governance should focus on designing legal, technical, and organizational structures that constrain and render machine behavior auditable and accountable, while keeping moral responsibility firmly anchored in human agents and institutions.

Philosophy & TechnologyVol. 39(4)
Openalex Percentile: Top 7%
Ethics and Social Impacts of AI
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.