MacroRisk-agent: a retrieval-augmented multi-agent framework with deterministic validation gates for auditable stress testing
Abstract Large language model (LLM)-based multi-agent systems face three open challenges in high-stakes decision support: hallucination control in structured output generation, coupling of semantic reasoning with deterministic numerical simulation, and end-to-end auditability of multi-step agentic workflows. We present MacroRisk-Agent , a multi-agent architecture that addresses these challenges in the domain of financial stress testing—a setting that demands strict output validity, reproducibility, and regulatory traceability. MacroRisk-Agent introduces three technical contributions: (1) a retrieval-augmented generation (RAG) pipeline with deterministic validation gates that converts natural-language stress intent into executable, schema-constrained parameter objects, achieving 92.6% validity rate versus 78.4% for RAG-augmented baselines ( $$p<0.001$$ ); (2) a constraint-coupled multi-agent simulation that models endogenous amplification through fire-sale, margin, and liquidity feedback loops, producing a 52% higher systemic-loss estimate than the one-pass configuration under the calibrated experimental setting; and (3) a counterfactual re-simulation framework that validates proposed interventions under identical scenario constraints, achieving 68.6% simulation-defined top-1 decision success rate. We evaluate MacroRisk-Agent as a proof-of-concept on public financial datasets (FRED, Yahoo Finance), a library of 150 curated crisis events, and 10 synthetic institutional portfolios; reported improvements are statistically significant across scenario validity, propagation modeling, and intervention quality within this synthetic evaluation environment . We position the framework as a proposed architecture—rather than an empirically validated, deployment-ready system—for building hallucination-controlled, provenance-tracked LLM-agent pipelines that are designed to improve auditability in high-stakes settings; validation on real institutional data is left to future work.
Authors
- Junyi Shangguan (ORCID: https://orcid.org/0009-0006-1459-6978)
- Qianyi Wang (ORCID: https://orcid.org/0000-0002-3679-4792)
- Yan Liu
- Shangfei Xu
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-10-08
- DOI
- https://doi.org/10.1038/s41598-026-73453-3
- Primary Topic
- Banking stability, regulation, efficiency
- Type
- article
- Field-Weighted Citation Impact
- 0.00