Document-Level Claim Extraction and Decontextualization Method Based on Large Language Models
Claim extraction and decontextualization are fundamental steps in automated fact-checking tasks. Claim extraction aims to identify core check-worthy claims from unverified text, while decontextualization seeks to eliminate the contextual dependencies of core claims so that they can be understood independently of the original text. In practical applications, information to be verified is usually presented in the form of documents, and existing document-level claim extraction and decontextualization methods often rely on complex multi-model architectures. This paper proposes LLM-CED, a document-level claim extraction and decontextualization method based on large language models, which consists of four modules: core-claim extraction, ambiguous-unit identification and question generation, question-answering-driven decontextualization, and self-reflective review. By using a single general-purpose LLM as the shared underlying model across all stages, LLM-CED reduces model heterogeneity and simplifies the integration and coordination of different processing stages. Experiments on the AVeriTeC-DCE dataset show that LLM-CED outperforms the evaluated baseline methods in document-level claim extraction, decontextualization, and the overall task. For the overall document-level claim extraction and decontextualization task, LLM-CED achieves a chrF score of 33.52% under the top-three evaluation setting.
Authors
- Tao He (ORCID: https://orcid.org/0000-0002-1074-2089)
- Meini Yang
- Zhihong Sun
- Renkang Hong
- Wei Hu
Institutions
- Naval University of Engineering (CN)
Publication Details
- Journal
- Applied Sciences
- Published
- 2026-09-24
- DOI
- https://doi.org/10.3390/app16199496
- Primary Topic
- Topic Modeling
- Type
- article
- Field-Weighted Citation Impact
- 0.00