Moral Reasoning RAG Evaluation Study: Does a Structured Moral-Reasoning RAG Improve Rubric- Scored Moral Reasoning in LLM Responses? A Paired, Blinded, Multi-Judge Evaluation

This record contains the research paper reporting a controlled evaluation of retrieval-augmented generation (RAG) designed to improve moral reasoning in large language models (LLMs). The underlying dataset, reproducibility materials, statistical analyses, and supporting files are archived separately and are available at: https://zenodo.org/records/22756832 The study examines whether a Moral Reasoning RAG system can improve the quality of ethical reasoning produced by generative artificial intelligence compared with a raw large language model API. The RAG condition incorporated structured retrieval, a moral reasoning ontology, and a critical reasoning directive intended to encourage more explicit consideration of human values, competing interests, consequences, duties, fairness, autonomy, human agency, and ethical principles. A total of 300 moral-reasoning questions were evaluated under two conditions: a Raw API condition and a Moral Reasoning RAG condition. Responses were assessed using blinded evaluation by three independent AI judge families across a 12-dimension moral reasoning rubric. The RAG condition achieved a higher mean Moral Reasoning Total Score than the Raw API condition (48.93 vs. 46.46), with a mean paired improvement of 2.48 points, 95% CI [1.91, 3.04], t(299) = 8.59, p = 4.86 × 10⁻¹⁶, and Cohen’s dz = 0.50. All 12 moral-reasoning dimensions improved under the RAG condition. In blinded pairwise comparisons, RAG responses were preferred in 55.56% of all evaluations and in 71.33% of decisive comparisons. The paper contributes empirical evidence to research on AI ethics, artificial intelligence ethics, ethical AI, responsible AI, trustworthy AI, moral AI, machine ethics, AI moral reasoning, AI ethical reasoning, automated ethical reasoning, machine morality, artificial moral agents, ethical decision-making, moral decision-making, AI alignment, value alignment, human values, human-centered AI, human agency, human oversight, AI accountability, AI transparency, AI explainability, AI fairness, algorithmic fairness, AI bias, social responsibility, human flourishing, AI governance, ethical AI governance, responsible AI governance, AI safety, AI risk management, generative AI ethics, large language model ethics, LLM ethics, LLM alignment, responsible LLM development, ethical LLM design, retrieval-augmented generation, and AI evaluation. The findings are relevant to the broader challenge of aligning artificial intelligence systems with human values while preserving transparency, accountability, autonomy, fairness, and meaningful human oversight. Rather than treating AI ethics solely as a matter of policy, regulation, or post-deployment governance, the study investigates whether moral and ethical reasoning can be strengthened directly at the response-generation stage through retrieval-augmented generation and structured reasoning support. This record contains the paper only. Researchers seeking the study dataset, reproducibility package, statistical outputs, evaluation materials, and supporting documentation should use the associated research archive: https://zenodo.org/records/22756832

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-29
DOI
https://doi.org/10.5281/zenodo.23029987
Primary Topic
Ethics and Social Impacts of AI
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Moral Reasoning RAG Evaluation Study: Does a Structured Moral-Reasoning RAG Improve Rubric- Scored Moral Reasoning in LLM Responses? A Paired, Blinded, Multi-Judge Evaluation

Ryan Campbell
Zenodo (CERN European Organization for Nuclear Research)
Ethics and Social Impacts of AI
preprint

Moral Reasoning RAG Evaluation Study: Does a Structured Moral-Reasoning RAG Improve Rubric- Scored Moral Reasoning in LLM Responses? A Paired, Blinded, Multi-Judge Evaluation

Ryan Campbell
preprint en

Abstract

This record contains the research paper reporting a controlled evaluation of retrieval-augmented generation (RAG) designed to improve moral reasoning in large language models (LLMs). The underlying dataset, reproducibility materials, statistical analyses, and supporting files are archived separately and are available at: https://zenodo.org/records/22756832 The study examines whether a Moral Reasoning RAG system can improve the quality of ethical reasoning produced by generative artificial intelligence compared with a raw large language model API. The RAG condition incorporated structured retrieval, a moral reasoning ontology, and a critical reasoning directive intended to encourage more explicit consideration of human values, competing interests, consequences, duties, fairness, autonomy, human agency, and ethical principles. A total of 300 moral-reasoning questions were evaluated under two conditions: a Raw API condition and a Moral Reasoning RAG condition. Responses were assessed using blinded evaluation by three independent AI judge families across a 12-dimension moral reasoning rubric. The RAG condition achieved a higher mean Moral Reasoning Total Score than the Raw API condition (48.93 vs. 46.46), with a mean paired improvement of 2.48 points, 95% CI [1.91, 3.04], t(299) = 8.59, p = 4.86 × 10⁻¹⁶, and Cohen’s dz = 0.50. All 12 moral-reasoning dimensions improved under the RAG condition. In blinded pairwise comparisons, RAG responses were preferred in 55.56% of all evaluations and in 71.33% of decisive comparisons. The paper contributes empirical evidence to research on AI ethics, artificial intelligence ethics, ethical AI, responsible AI, trustworthy AI, moral AI, machine ethics, AI moral reasoning, AI ethical reasoning, automated ethical reasoning, machine morality, artificial moral agents, ethical decision-making, moral decision-making, AI alignment, value alignment, human values, human-centered AI, human agency, human oversight, AI accountability, AI transparency, AI explainability, AI fairness, algorithmic fairness, AI bias, social responsibility, human flourishing, AI governance, ethical AI governance, responsible AI governance, AI safety, AI risk management, generative AI ethics, large language model ethics, LLM ethics, LLM alignment, responsible LLM development, ethical LLM design, retrieval-augmented generation, and AI evaluation. The findings are relevant to the broader challenge of aligning artificial intelligence systems with human values while preserving transparency, accountability, autonomy, fairness, and meaningful human oversight. Rather than treating AI ethics solely as a matter of policy, regulation, or post-deployment governance, the study investigates whether moral and ethical reasoning can be strengthened directly at the response-generation stage through retrieval-augmented generation and structured reasoning support. This record contains the paper only. Researchers seeking the study dataset, reproducibility package, statistical outputs, evaluation materials, and supporting documentation should use the associated research archive: https://zenodo.org/records/22756832

Zenodo (CERN European Organization for Nuclear Research)
Peace, Justice and strong institutions
Ethics and Social Impacts of AI
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.