Selective Invariance Violations in Large Language Model Moral Judgment: A Geometric Framework for Behavioral Manipulation Detection
Large language models are increasingly deployed in safety-critical decision services---content moderation, clinical decision support, legal analysis---yet methods for characterizing their vulnerability profiles across multiple attack surfaces remain underdeveloped. We introduce a geometric evaluation framework that maps moral judgment to a 7-dimensional harm space, applies five qualitatively distinct perturbation types across five cognitive domains, and produces per-model vulnerability profiles that reveal manipulations each model resists and it does not. Testing 5~models spanning 2~architecture families under a realistic \\50/day compute budget, we find that vulnerabilities are : linguistic framing, emotional anchoring, and irrelevant sensory detail reliably displace judgments, while gender swap and evaluation order do not---identifying salience manipulation as the specific attack surface. The framework further reveals that robustness profiles are partially dissociable across models: a model with zero sycophancy has the worst emotional anchoring recovery; a model with the best anchoring recovery has the worst working memory. No single robustness score captures these structures. An inter-model agreement study over an independent open-model panel confirms these harm dimensions are reliably measurable (ICC(2,k)\\,=\\,0.97), and a worked content-moderation exploit shows that salience manipulation flips not only a scalar harm threshold but the typed verdict of a downstream rule-based decision kernel. The evaluation pipeline---with adaptive concurrency, budget-aware execution, and three-tier data---scales to new models and perturbation types within fixed compute constraints, providing a practical tool for multi-dimensional LLM security assessment. Target venue: IEEE BigDataService 2026 (accepted). Author preprint deposited for archival and citation. Draft — pending author review.
Authors
- Andrew Bond
Institutions
- San Jose State University (US)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-07-19
- DOI
- https://doi.org/10.5281/zenodo.21435283
- Primary Topic
- Adversarial Robustness in Machine Learning
- Type
- article
- Field-Weighted Citation Impact
- 0.00