Grounding language models with deterministic verifiers for fraud detection

Abstract Financial fraud detection is complicated by severe class imbalance, temporal change, adversarial manipulation, and the need for evidence-grounded explanations. Although large language models can generate readable rationales, their claims may be unsupported when they are not independently checked against structured evidence. We propose VR-FraudNet, a five-stage framework combining a time-conditioned spectral graph encoder, LightGBM triage and threshold routing, a schema-constrained rationale language model, a fixed deterministic verifier, and an isotonic probability mixer with split-conformal calibration. The language model proposes machine-checkable claims for review-band transactions, while the verifier evaluates each claim against the available transaction, temporal, and graph evidence. Verifier-rejected rationales are excluded from automated use, and the corresponding transactions are assigned to human-review escalation. D1 BAF, D2 AMLworld HI-Small, and D3 IEEE-CIS are used as the primary predictive benchmarks, whereas D4 Elliptic++ and D5 DGraph-Fin support complementary temporal-graph, strict-inductive, temporal-robustness, and message-passing stress tests. VR-FraudNet improves AUPRC over TabTransformer from 0.5012 to 0.5247 on D1, from 0.5137 to 0.5824 on D2, and from 0.7234 to 0.7456 on D3. The predeclared 5-percentage-point temporal-degradation target is satisfied in the evaluated D1 month-drift, D3 late-fold, and D4 pre-shock comparisons. However, D4 post-shock AUPRC decreases from 0.6823 to 0.6087, an absolute reduction of 7.36 percentage points that exceeds the target. Under the four-family structured predictive edit set, mean evasion rates are 0.0667, 0.1016, and 0.1672 for budgets 1, 2, and 4, respectively. Under the stated operational cost assumptions, VR-FraudNet produces lower total loss than TabTransformer, with the largest difference on D2. Measured p99 latency is 10.23 ms on the common path and 175.23 ms on the escalation path under the reported hardware configuration. Empirical split-conformal miscoverage remains below the stated bounds in the reported calibration settings, while the formal marginal-coverage statement applies only under the stated exchangeability assumptions. These findings support verifier-grounded fraud detection under the evaluated public datasets and protocols but do not establish universal adversarial safety, temporal robustness, regulatory sufficiency, or deployment readiness.

Authors

Publication Details

Journal
Discover Artificial Intelligence
Published
2026-10-08
DOI
https://doi.org/10.1007/s44163-026-02355-0
Primary Topic
Imbalanced Data Classification Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Grounding language models with deterministic verifiers for fraud detection

Partha Chakraborty, Abdul Basit, Md Sultanul Arefin Sourav, Evha Rozario et al.
Discover Artificial Intelligence
Imbalanced Data Classification Techniques
article

Grounding language models with deterministic verifiers for fraud detection

Partha Chakraborty, Abdul Basit, Md Sultanul Arefin Sourav, Evha Rozario, Kaniz Sultana Chy, Md Ashiqul Islam, Samia Akter, Maria Kabtia, Mohammad Sazzad Hossain, Md Shahiduzzaman
article en

Abstract

Abstract Financial fraud detection is complicated by severe class imbalance, temporal change, adversarial manipulation, and the need for evidence-grounded explanations. Although large language models can generate readable rationales, their claims may be unsupported when they are not independently checked against structured evidence. We propose VR-FraudNet, a five-stage framework combining a time-conditioned spectral graph encoder, LightGBM triage and threshold routing, a schema-constrained rationale language model, a fixed deterministic verifier, and an isotonic probability mixer with split-conformal calibration. The language model proposes machine-checkable claims for review-band transactions, while the verifier evaluates each claim against the available transaction, temporal, and graph evidence. Verifier-rejected rationales are excluded from automated use, and the corresponding transactions are assigned to human-review escalation. D1 BAF, D2 AMLworld HI-Small, and D3 IEEE-CIS are used as the primary predictive benchmarks, whereas D4 Elliptic++ and D5 DGraph-Fin support complementary temporal-graph, strict-inductive, temporal-robustness, and message-passing stress tests. VR-FraudNet improves AUPRC over TabTransformer from 0.5012 to 0.5247 on D1, from 0.5137 to 0.5824 on D2, and from 0.7234 to 0.7456 on D3. The predeclared 5-percentage-point temporal-degradation target is satisfied in the evaluated D1 month-drift, D3 late-fold, and D4 pre-shock comparisons. However, D4 post-shock AUPRC decreases from 0.6823 to 0.6087, an absolute reduction of 7.36 percentage points that exceeds the target. Under the four-family structured predictive edit set, mean evasion rates are 0.0667, 0.1016, and 0.1672 for budgets 1, 2, and 4, respectively. Under the stated operational cost assumptions, VR-FraudNet produces lower total loss than TabTransformer, with the largest difference on D2. Measured p99 latency is 10.23 ms on the common path and 175.23 ms on the escalation path under the reported hardware configuration. Empirical split-conformal miscoverage remains below the stated bounds in the reported calibration settings, while the formal marginal-coverage statement applies only under the stated exchangeability assumptions. These findings support verifier-grounded fraud detection under the evaluated public datasets and protocols but do not establish universal adversarial safety, temporal robustness, regulatory sufficiency, or deployment readiness.

Discover Artificial IntelligenceVol. 6(1)
Openalex Percentile: Top 12%
Imbalanced Data Classification Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.