Large language models for literature screening in conceptually complex and interdisciplinary reviews
Purpose This study examines the capabilities and limitations of large language models (LLMs) for literature screening in systematic reviews involving conceptually diffuse and interdisciplinary topics. Design/methodology/approach Two screening datasets were constructed. Four LLMs, Claude 4.6, DeepSeek V4 Pro, Gemini 3.1 Pro, and GPT-5.5, were evaluated using a staged, rule-guided prompting workflow covering title screening, abstract screening, and full-text assessment. Model performance was assessed using accuracy, precision, recall, specificity, the F1-score, and Cohen's kappa, supplemented by stage-wise screening flow analysis and qualitative discrepancy analysis between model and human reviewer decisions. Findings Model performance varied substantially across datasets and models. In the conceptually diffuse Dataset A, all models achieved high specificity, but positive-class performance was more limited. DeepSeek V4 Pro achieved the highest accuracy, precision, F1-score, and Cohen's kappa, whereas Claude 4.6 achieved the highest recall. In the more clearly bounded Dataset B, recall was higher and model behavior was more convergent, with GPT-5.5 showing the best overall balance across performance indicators. Stage-wise analysis showed that models differed in where they made inclusion and exclusion decisions, while discrepancy analysis indicated that errors were mainly related to conceptual scope confusion, context misidentification, and underestimation of analytical depth. These findings suggest that current LLMs can support literature screening, but their reliability depends strongly on the conceptual clarity of the review task and the structure of eligibility criteria. Originality/value This study extends the literature on LLM-assisted screening by focusing on conceptually abstract and interdisciplinary review tasks. This study introduces a staged, Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA)-aligned screening workflow and conducts an error analysis that explains how LLM screening can fail in semantically diffuse settings. The study provides practical implications for the transparent and responsible use of LLMs in informetric research and systematic review practice.
Authors
- Xuan Han (ORCID: https://orcid.org/0000-0002-4685-8235)
- Ni Cheng (ORCID: https://orcid.org/0000-0001-7058-6180)
- Heng Dong (ORCID: https://orcid.org/0000-0002-7736-4940)
- Shuzhen Zhu
- Beibei Tan
Institutions
- Huazhong Agricultural University (CN)
- EA Technology (GB)
- Hubei Provincial Center for Disease Control and Prevention (CN)
- Zhejiang University (CN)
Publication Details
- Journal
- Aslib Journal of Information Management
- Published
- 2026-09-17
- DOI
- https://doi.org/10.1108/ajim-06-2025-0363
- Primary Topic
- Meta-analysis and systematic reviews
- Type
- article
- Field-Weighted Citation Impact
- 0.00