Document-Level Claim Extraction and Decontextualization Method Based on Large Language Models

Claim extraction and decontextualization are fundamental steps in automated fact-checking tasks. Claim extraction aims to identify core check-worthy claims from unverified text, while decontextualization seeks to eliminate the contextual dependencies of core claims so that they can be understood independently of the original text. In practical applications, information to be verified is usually presented in the form of documents, and existing document-level claim extraction and decontextualization methods often rely on complex multi-model architectures. This paper proposes LLM-CED, a document-level claim extraction and decontextualization method based on large language models, which consists of four modules: core-claim extraction, ambiguous-unit identification and question generation, question-answering-driven decontextualization, and self-reflective review. By using a single general-purpose LLM as the shared underlying model across all stages, LLM-CED reduces model heterogeneity and simplifies the integration and coordination of different processing stages. Experiments on the AVeriTeC-DCE dataset show that LLM-CED outperforms the evaluated baseline methods in document-level claim extraction, decontextualization, and the overall task. For the overall document-level claim extraction and decontextualization task, LLM-CED achieves a chrF score of 33.52% under the top-three evaluation setting.

Authors

Institutions

Publication Details

Journal
Applied Sciences
Published
2026-09-24
DOI
https://doi.org/10.3390/app16199496
Primary Topic
Topic Modeling
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Document-Level Claim Extraction and Decontextualization Method Based on Large Language Models

Tao He, Meini Yang, Zhihong Sun, Renkang Hong et al.
Applied Sciences
Topic Modeling
article

Document-Level Claim Extraction and Decontextualization Method Based on Large Language Models

Tao He, Meini Yang, Zhihong Sun, Renkang Hong, Wei Hu
article en

Abstract

Claim extraction and decontextualization are fundamental steps in automated fact-checking tasks. Claim extraction aims to identify core check-worthy claims from unverified text, while decontextualization seeks to eliminate the contextual dependencies of core claims so that they can be understood independently of the original text. In practical applications, information to be verified is usually presented in the form of documents, and existing document-level claim extraction and decontextualization methods often rely on complex multi-model architectures. This paper proposes LLM-CED, a document-level claim extraction and decontextualization method based on large language models, which consists of four modules: core-claim extraction, ambiguous-unit identification and question generation, question-answering-driven decontextualization, and self-reflective review. By using a single general-purpose LLM as the shared underlying model across all stages, LLM-CED reduces model heterogeneity and simplifies the integration and coordination of different processing stages. Experiments on the AVeriTeC-DCE dataset show that LLM-CED outperforms the evaluated baseline methods in document-level claim extraction, decontextualization, and the overall task. For the overall document-level claim extraction and decontextualization task, LLM-CED achieves a chrF score of 33.52% under the top-three evaluation setting.

Applied SciencesVol. 16(19)
Naval University of Engineering (CN)
Openalex Percentile: Top 9%
Topic Modeling
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Document-Level Claim Extraction and Decontextualization Method Based on Large Language Models — Tao He, Meini Yang, et al. · Applied Sciences (2026) | TGRS Research Map | TGRS