Countering the Risks of AI Overreliance: Leveraging Multi-Agent Debate for Complementary Human–AI Decision-Making in Legal Contexts
Despite their promise, human–AI teams often fail to outperform either humans or AI alone, in part due to overreliance—where users defer to AI recommendations even when they are incorrect. Prior mitigation strategies, such as explainable AI and trust calibration, have shown limited effectiveness, as they largely preserve a single-recommendation paradigm that anchors user judgment. We introduce ARMADA (Argumentation and Resolution Multi-Agent Debate Architecture), a human-centered approach that restructures AI assistance as a multi-agent debate. Instead of producing a single recommendation, multiple AI agents present competing arguments, leaving the human to make the final decision. We evaluate ARMADA in a controlled user study (N = 186) using civil legal reasoning tasks as a testbed. Results show that ARMADA significantly improves team accuracy by 12% in high-difficulty conditions where AI performance is unreliable. These findings indicate that debate-based interaction can reduce overreliance and enable complementary human–AI performance under specific conditions, though benefits are not uniform across all settings. Importantly, ARMADA does not improve the underlying correctness of AI outputs but instead changes how users engage with them. We discuss the conditions under which complementarity emerges, along with implications for designing human-centered AI-assisted decision-making systems.
Authors
- Brydon T. Wang (ORCID: https://orcid.org/0000-0002-1975-2689)
- Tim Miller (ORCID: https://orcid.org/0000-0003-4908-6063)
- Gianluca Demartini (ORCID: https://orcid.org/0000-0002-7311-3693)
- Bing Tuo (ORCID: https://orcid.org/0009-0001-4123-5346)
Institutions
- The University of Queensland (AU)
Publication Details
- Journal
- ACM Transactions on Intelligent Systems and Technology
- Published
- 2026-09-17
- DOI
- https://doi.org/10.1145/3848504
- Primary Topic
- Ethics and Social Impacts of AI
- Type
- article
- Field-Weighted Citation Impact
- 0.00