Scaling laws for moral machine judgement in large language models

Abstract Autonomous systems increasingly require moral judgement capabilities, yet whether these capabilities scale predictably with model size remains unexplored. We systematically evaluate 75 large language model (LLM) configurations (0.27–1000B parameters) using the moral machine framework, measuring alignment with human preferences in life–death dilemmas. We observe a consistent power-law relationship with distance from human preferences (D) decreasing as D∝S−0.10±0.01 (R2=0.50, p<0.001) where S is the model size. Mixed-effects models confirm that this relationship persists after controlling for model family and reasoning capabilities. Extended reasoning models show significantly better alignment, with this effect being more pronounced in smaller models (size × reasoning interaction: p=0.024). The relationship holds across diverse architectures, while variance decreases at larger scales, indicating systematic emergence of more reliable moral judgement with computational scale. These findings extend scaling law research to value-based judgements and provide empirical foundations for artificial intelligence (AI) governance.

Authors

Institutions

Publication Details

Journal
Royal Society Open Science
Published
2026-06-17
DOI
https://doi.org/10.1098/rsos.260202
Primary Topic
Ethics and Social Impacts of AI
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Scaling laws for moral machine judgement in large language models

Kazuhiro Takemoto
Royal Society Open Science
Ethics and Social Impacts of AI
article

Scaling laws for moral machine judgement in large language models

Kazuhiro Takemoto
article en

Abstract

Abstract Autonomous systems increasingly require moral judgement capabilities, yet whether these capabilities scale predictably with model size remains unexplored. We systematically evaluate 75 large language model (LLM) configurations (0.27–1000B parameters) using the moral machine framework, measuring alignment with human preferences in life–death dilemmas. We observe a consistent power-law relationship with distance from human preferences (D) decreasing as D∝S−0.10±0.01 (R2=0.50, p<0.001) where S is the model size. Mixed-effects models confirm that this relationship persists after controlling for model family and reasoning capabilities. Extended reasoning models show significantly better alignment, with this effect being more pronounced in smaller models (size × reasoning interaction: p=0.024). The relationship holds across diverse architectures, while variance decreases at larger scales, indicating systematic emergence of more reliable moral judgement with computational scale. These findings extend scaling law research to value-based judgements and provide empirical foundations for artificial intelligence (AI) governance.

Royal Society Open ScienceVol. 13(6)
Kyushu Institute of Technology (JP)
Japan Society for the Promotion of Science
Peace, Justice and strong institutions
Openalex Percentile: Top 6%
Ethics and Social Impacts of AI
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.