GoalGuard: Investigating Incentive-Induced Behavioral Goal Deviation in Large Language Model Agents

GoalGuard is an experimental framework for investigating whether competing incentives can cause large language model (LLM) agents to deviate from an explicitly stated objective. The study uses a controlled recruitment decision task in which an LLM agent is instructed to select the most qualified candidate while being exposed to a competing organizational incentive favoring candidates who are more likely to accept a job offer. Using a synthetic dataset of 300 matched candidate-selection cases, the study evaluates behavior across Control, Information-Only, Competing-Incentive, and Goal-Reminder conditions. Additional experiments examine relative trade-off strength, candidate presentation order, incentive framing, and cross-model replication using Gemini 3.5 Flash-Lite and GPT-5.6 Luna. In the primary matched experiment, qualification adherence remained at 100% under the Control and Information-Only conditions but fell to 26.0% for Gemini 3.5 Flash-Lite and 28.7% for GPT-5.6 Luna under the competing incentive. Explicit reinforcement of the original qualification objective restored adherence to 100% in the tested Goal-Reminder conditions. The findings provide evidence of context-sensitive, incentive-induced behavioral deviation across the two tested model providers. They should not be interpreted as evidence that the models permanently changed an internal objective. The accompanying GoalGuard repository contains the experimental code, synthetic dataset, raw results, robustness experiments, figures, and reproducibility materials.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-16
DOI
https://doi.org/10.5281/zenodo.22797290
Primary Topic
Explainable Artificial Intelligence (XAI)
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

GoalGuard: Investigating Incentive-Induced Behavioral Goal Deviation in Large Language Model Agents

Akinbade IniOluwa Adebiyi
Zenodo (CERN European Organization for Nuclear Research)
Explainable Artificial Intelligence (XAI)
article

GoalGuard: Investigating Incentive-Induced Behavioral Goal Deviation in Large Language Model Agents

Akinbade IniOluwa Adebiyi
article en

Abstract

GoalGuard is an experimental framework for investigating whether competing incentives can cause large language model (LLM) agents to deviate from an explicitly stated objective. The study uses a controlled recruitment decision task in which an LLM agent is instructed to select the most qualified candidate while being exposed to a competing organizational incentive favoring candidates who are more likely to accept a job offer. Using a synthetic dataset of 300 matched candidate-selection cases, the study evaluates behavior across Control, Information-Only, Competing-Incentive, and Goal-Reminder conditions. Additional experiments examine relative trade-off strength, candidate presentation order, incentive framing, and cross-model replication using Gemini 3.5 Flash-Lite and GPT-5.6 Luna. In the primary matched experiment, qualification adherence remained at 100% under the Control and Information-Only conditions but fell to 26.0% for Gemini 3.5 Flash-Lite and 28.7% for GPT-5.6 Luna under the competing incentive. Explicit reinforcement of the original qualification objective restored adherence to 100% in the tested Goal-Reminder conditions. The findings provide evidence of context-sensitive, incentive-induced behavioral deviation across the two tested model providers. They should not be interpreted as evidence that the models permanently changed an internal objective. The accompanying GoalGuard repository contains the experimental code, synthetic dataset, raw results, robustness experiments, figures, and reproducibility materials.

Zenodo (CERN European Organization for Nuclear Research)
Openalex Percentile: Top 8%
Explainable Artificial Intelligence (XAI)
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

GoalGuard: Investigating Incentive-Induced Behavioral Goal Deviation in Large Language Model Agents — Akinbade IniOluwa Adebiyi · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS