When Does Claim Decomposition Help? Fine-Grained vs. Single-Prompt LLM Judges for RAG Hallucination Detection
Retrieval-Augmented Generation (RAG) systems are commonly evaluated for hallucination using either fine-grained, claim-decomposition-based checkers (e.g., RAGChecker, FActScore) or coarse, single-prompt LLM-as-judge baselines. A common assumption in the RAG evaluation literature is that decomposition-based checking is more reliable, since it can assign partial credit and diagnose why an answer is wrong rather than only whether it is wrong. I built a RAGChecker-style checker from scratch, with atomic claim extraction, per-claim entailment checking against retrieved context, and eleven diagnostic metrics spanning retrieval and generation quality. I evaluated it against a single-prompt LLM judge on two datasets: a small, topically well-separated synthetic knowledge base (14 company profiles), and a harder real-world knowledge base of 14 historical and scientific topics that deliberately includes confusable pairs (e.g., two mountains, two physicists, two inventors). Ground truth was established by deliberately corrupting a known subset of generated answers. On the easy dataset, both checkers achieved perfect classification (F1 = 1.000). On the harder dataset, the fine-grained checker's F1 fell to 0.714 while the single-prompt judge remained perfect (F1 = 1.000). Error analysis points to two failure modes of claim-decomposition checking: (1) claim-extraction and reference-entailment errors that let wrong answers through, producing false negatives (in one case the extractor rewrote a false assertion as a true negation, and in another the gold-answer judge credited year and location details that the gold never states); and (2) phrasing mismatches between extracted claims and single-sentence gold answers on multi-hop comparison questions, producing false positives. In this setting, decomposition-based checking was not superior to single-prompt judging, and its advantage appears to depend on claim-extraction quality and on how answers are matched to a reference. The study uses one small model and small samples, so the findings are a demonstration of the failure modes rather than a general ranking of methods.
Authors
- Mrunal Gangurde
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-29
- DOI
- https://doi.org/10.5281/zenodo.23038176
- Primary Topic
- Misinformation and Its Impacts
- Type
- preprint