Selective Invariance Violations in Large Language Model Moral Judgment: A Geometric Framework for Behavioral Manipulation Detection

Large language models are increasingly deployed in safety-critical decision services---content moderation, clinical decision support, legal analysis---yet methods for characterizing their vulnerability profiles across multiple attack surfaces remain underdeveloped. We introduce a geometric evaluation framework that maps moral judgment to a 7-dimensional harm space, applies five qualitatively distinct perturbation types across five cognitive domains, and produces per-model vulnerability profiles that reveal manipulations each model resists and it does not. Testing 5~models spanning 2~architecture families under a realistic \\50/day compute budget, we find that vulnerabilities are : linguistic framing, emotional anchoring, and irrelevant sensory detail reliably displace judgments, while gender swap and evaluation order do not---identifying salience manipulation as the specific attack surface. The framework further reveals that robustness profiles are partially dissociable across models: a model with zero sycophancy has the worst emotional anchoring recovery; a model with the best anchoring recovery has the worst working memory. No single robustness score captures these structures. An inter-model agreement study over an independent open-model panel confirms these harm dimensions are reliably measurable (ICC(2,k)\\,=\\,0.97), and a worked content-moderation exploit shows that salience manipulation flips not only a scalar harm threshold but the typed verdict of a downstream rule-based decision kernel. The evaluation pipeline---with adaptive concurrency, budget-aware execution, and three-tier data---scales to new models and perturbation types within fixed compute constraints, providing a practical tool for multi-dimensional LLM security assessment. Target venue: IEEE BigDataService 2026 (accepted). Author preprint deposited for archival and citation. Draft — pending author review.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-07-19
DOI
https://doi.org/10.5281/zenodo.21435283
Primary Topic
Adversarial Robustness in Machine Learning
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Selective Invariance Violations in Large Language Model Moral Judgment: A Geometric Framework for Behavioral Manipulation Detection

Andrew Bond
Zenodo (CERN European Organization for Nuclear Research)
Adversarial Robustness in Machine Learning
article

Selective Invariance Violations in Large Language Model Moral Judgment: A Geometric Framework for Behavioral Manipulation Detection

Andrew Bond
article en

Abstract

Large language models are increasingly deployed in safety-critical decision services---content moderation, clinical decision support, legal analysis---yet methods for characterizing their vulnerability profiles across multiple attack surfaces remain underdeveloped. We introduce a geometric evaluation framework that maps moral judgment to a 7-dimensional harm space, applies five qualitatively distinct perturbation types across five cognitive domains, and produces per-model vulnerability profiles that reveal manipulations each model resists and it does not. Testing 5~models spanning 2~architecture families under a realistic \50/day compute budget, we find that vulnerabilities are : linguistic framing, emotional anchoring, and irrelevant sensory detail reliably displace judgments, while gender swap and evaluation order do not---identifying salience manipulation as the specific attack surface. The framework further reveals that robustness profiles are partially dissociable across models: a model with zero sycophancy has the worst emotional anchoring recovery; a model with the best anchoring recovery has the worst working memory. No single robustness score captures these structures. An inter-model agreement study over an independent open-model panel confirms these harm dimensions are reliably measurable (ICC(2,k)\,=\,0.97), and a worked content-moderation exploit shows that salience manipulation flips not only a scalar harm threshold but the typed verdict of a downstream rule-based decision kernel. The evaluation pipeline---with adaptive concurrency, budget-aware execution, and three-tier data---scales to new models and perturbation types within fixed compute constraints, providing a practical tool for multi-dimensional LLM security assessment. Target venue: IEEE BigDataService 2026 (accepted). Author preprint deposited for archival and citation. Draft — pending author review.

Zenodo (CERN European Organization for Nuclear Research)
San Jose State University (US)
Gender equality
Openalex Percentile: Top 6%
Adversarial Robustness in Machine Learning
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.