Mitigating Hallucinations in Finance-Based Multi-Agent Large Language Model Systems

Large language models are increasingly used for financial question answering, while they are prone to generating hallucinated content. In this research, we propose a multi-signal framework for hallucination detection and mitigation. Our framework combines six signals (entailment, semantic similarity, claim verification, numeric consistency, token overlap and confidence calibration techniques) to identify hallucination type and severity. We evaluated our framework using multiple language models across three financial QA benchmark datasets (FinQA, FinanceQA, and FinDER). Our proposed framework corrects up to 78.4% of the detected hallucinations (fix rate on the hallucinated subset of FinQA; 47.7% on FinanceQA and 45.7% on FinDER), outperforming a type-agnostic mitigation baseline on all three datasets. This highlights that targeted, type-aware hallucination mitigation can significantly improve answer reliability while remaining computationally effective.

Authors

Institutions

Publication Details

Journal
Electronics
Published
2026-09-28
DOI
https://doi.org/10.3390/electronics15194456
Primary Topic
Topic Modeling
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Mitigating Hallucinations in Finance-Based Multi-Agent Large Language Model Systems

Rashmi Nagpal, Amar Deep Gupta, Unyimeabasi Usua, Kailey Simons et al.
Electronics
Topic Modeling
article

Mitigating Hallucinations in Finance-Based Multi-Agent Large Language Model Systems

Rashmi Nagpal, Amar Deep Gupta, Unyimeabasi Usua, Kailey Simons, Vishal Gossain, Sabrina Queipo
article en

Abstract

Large language models are increasingly used for financial question answering, while they are prone to generating hallucinated content. In this research, we propose a multi-signal framework for hallucination detection and mitigation. Our framework combines six signals (entailment, semantic similarity, claim verification, numeric consistency, token overlap and confidence calibration techniques) to identify hallucination type and severity. We evaluated our framework using multiple language models across three financial QA benchmark datasets (FinQA, FinanceQA, and FinDER). Our proposed framework corrects up to 78.4% of the detected hallucinations (fix rate on the hallucinated subset of FinQA; 47.7% on FinanceQA and 45.7% on FinDER), outperforming a type-agnostic mitigation baseline on all three datasets. This highlights that targeted, type-aware hallucination mitigation can significantly improve answer reliability while remaining computationally effective.

ElectronicsVol. 15(19)
The University of Texas at El Paso (US), Massachusetts Institute of Technology (US)
Openalex Percentile: Top 9%
Topic Modeling
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Mitigating Hallucinations in Finance-Based Multi-Agent Large Language Model Systems — Rashmi Nagpal, Amar Deep Gupta, et al. · Electronics (2026) | TGRS Research Map | TGRS