Mitigating Hallucinations in Finance-Based Multi-Agent Large Language Model Systems
Large language models are increasingly used for financial question answering, while they are prone to generating hallucinated content. In this research, we propose a multi-signal framework for hallucination detection and mitigation. Our framework combines six signals (entailment, semantic similarity, claim verification, numeric consistency, token overlap and confidence calibration techniques) to identify hallucination type and severity. We evaluated our framework using multiple language models across three financial QA benchmark datasets (FinQA, FinanceQA, and FinDER). Our proposed framework corrects up to 78.4% of the detected hallucinations (fix rate on the hallucinated subset of FinQA; 47.7% on FinanceQA and 45.7% on FinDER), outperforming a type-agnostic mitigation baseline on all three datasets. This highlights that targeted, type-aware hallucination mitigation can significantly improve answer reliability while remaining computationally effective.
Authors
- Rashmi Nagpal (ORCID: https://orcid.org/0000-0002-8087-9806)
- Amar Deep Gupta (ORCID: https://orcid.org/0000-0001-9306-1256)
- Unyimeabasi Usua (ORCID: https://orcid.org/0009-0003-4387-0647)
- Kailey Simons
- Vishal Gossain
- Sabrina Queipo
Institutions
- The University of Texas at El Paso (US)
- Massachusetts Institute of Technology (US)
Publication Details
- Journal
- Electronics
- Published
- 2026-09-28
- DOI
- https://doi.org/10.3390/electronics15194456
- Primary Topic
- Topic Modeling
- Type
- article
- Field-Weighted Citation Impact
- 0.00