Countering the Risks of AI Overreliance: Leveraging Multi-Agent Debate for Complementary Human–AI Decision-Making in Legal Contexts

Despite their promise, human–AI teams often fail to outperform either humans or AI alone, in part due to overreliance—where users defer to AI recommendations even when they are incorrect. Prior mitigation strategies, such as explainable AI and trust calibration, have shown limited effectiveness, as they largely preserve a single-recommendation paradigm that anchors user judgment. We introduce ARMADA (Argumentation and Resolution Multi-Agent Debate Architecture), a human-centered approach that restructures AI assistance as a multi-agent debate. Instead of producing a single recommendation, multiple AI agents present competing arguments, leaving the human to make the final decision. We evaluate ARMADA in a controlled user study (N = 186) using civil legal reasoning tasks as a testbed. Results show that ARMADA significantly improves team accuracy by 12% in high-difficulty conditions where AI performance is unreliable. These findings indicate that debate-based interaction can reduce overreliance and enable complementary human–AI performance under specific conditions, though benefits are not uniform across all settings. Importantly, ARMADA does not improve the underlying correctness of AI outputs but instead changes how users engage with them. We discuss the conditions under which complementarity emerges, along with implications for designing human-centered AI-assisted decision-making systems.

Authors

Institutions

Publication Details

Journal
ACM Transactions on Intelligent Systems and Technology
Published
2026-09-17
DOI
https://doi.org/10.1145/3848504
Primary Topic
Ethics and Social Impacts of AI
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Countering the Risks of AI Overreliance: Leveraging Multi-Agent Debate for Complementary Human–AI Decision-Making in Legal Contexts

Brydon T. Wang, Tim Miller, Gianluca Demartini, Bing Tuo
ACM Transactions on Intelligent Systems and Technology
Ethics and Social Impacts of AI
article

Countering the Risks of AI Overreliance: Leveraging Multi-Agent Debate for Complementary Human–AI Decision-Making in Legal Contexts

Brydon T. Wang, Tim Miller, Gianluca Demartini, Bing Tuo
article en

Abstract

Despite their promise, human–AI teams often fail to outperform either humans or AI alone, in part due to overreliance—where users defer to AI recommendations even when they are incorrect. Prior mitigation strategies, such as explainable AI and trust calibration, have shown limited effectiveness, as they largely preserve a single-recommendation paradigm that anchors user judgment. We introduce ARMADA (Argumentation and Resolution Multi-Agent Debate Architecture), a human-centered approach that restructures AI assistance as a multi-agent debate. Instead of producing a single recommendation, multiple AI agents present competing arguments, leaving the human to make the final decision. We evaluate ARMADA in a controlled user study (N = 186) using civil legal reasoning tasks as a testbed. Results show that ARMADA significantly improves team accuracy by 12% in high-difficulty conditions where AI performance is unreliable. These findings indicate that debate-based interaction can reduce overreliance and enable complementary human–AI performance under specific conditions, though benefits are not uniform across all settings. Importantly, ARMADA does not improve the underlying correctness of AI outputs but instead changes how users engage with them. We discuss the conditions under which complementarity emerges, along with implications for designing human-centered AI-assisted decision-making systems.

ACM Transactions on Intelligent Systems and Technology
The University of Queensland (AU)
Peace, Justice and strong institutions
Openalex Percentile: Top 7%
Ethics and Social Impacts of AI
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Countering the Risks of AI Overreliance: Leveraging Multi-Agent Debate for Complementary Human–AI Decision-Making in Legal Contexts — Brydon T. Wang, Tim Miller, et al. · ACM Transactions on Intelligent Systems and Technology (2026) | TGRS Research Map | TGRS