Large language models for literature screening in conceptually complex and interdisciplinary reviews

Purpose This study examines the capabilities and limitations of large language models (LLMs) for literature screening in systematic reviews involving conceptually diffuse and interdisciplinary topics. Design/methodology/approach Two screening datasets were constructed. Four LLMs, Claude 4.6, DeepSeek V4 Pro, Gemini 3.1 Pro, and GPT-5.5, were evaluated using a staged, rule-guided prompting workflow covering title screening, abstract screening, and full-text assessment. Model performance was assessed using accuracy, precision, recall, specificity, the F1-score, and Cohen's kappa, supplemented by stage-wise screening flow analysis and qualitative discrepancy analysis between model and human reviewer decisions. Findings Model performance varied substantially across datasets and models. In the conceptually diffuse Dataset A, all models achieved high specificity, but positive-class performance was more limited. DeepSeek V4 Pro achieved the highest accuracy, precision, F1-score, and Cohen's kappa, whereas Claude 4.6 achieved the highest recall. In the more clearly bounded Dataset B, recall was higher and model behavior was more convergent, with GPT-5.5 showing the best overall balance across performance indicators. Stage-wise analysis showed that models differed in where they made inclusion and exclusion decisions, while discrepancy analysis indicated that errors were mainly related to conceptual scope confusion, context misidentification, and underestimation of analytical depth. These findings suggest that current LLMs can support literature screening, but their reliability depends strongly on the conceptual clarity of the review task and the structure of eligibility criteria. Originality/value This study extends the literature on LLM-assisted screening by focusing on conceptually abstract and interdisciplinary review tasks. This study introduces a staged, Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA)-aligned screening workflow and conducts an error analysis that explains how LLM screening can fail in semantically diffuse settings. The study provides practical implications for the transparent and responsible use of LLMs in informetric research and systematic review practice.

Authors

Institutions

Publication Details

Journal
Aslib Journal of Information Management
Published
2026-09-17
DOI
https://doi.org/10.1108/ajim-06-2025-0363
Primary Topic
Meta-analysis and systematic reviews
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Large language models for literature screening in conceptually complex and interdisciplinary reviews

Xuan Han, Ni Cheng, Heng Dong, Shuzhen Zhu et al.
Aslib Journal of Information Management
Meta-analysis and systematic reviews
article

Large language models for literature screening in conceptually complex and interdisciplinary reviews

Xuan Han, Ni Cheng, Heng Dong, Shuzhen Zhu, Beibei Tan
article en

Abstract

Purpose This study examines the capabilities and limitations of large language models (LLMs) for literature screening in systematic reviews involving conceptually diffuse and interdisciplinary topics. Design/methodology/approach Two screening datasets were constructed. Four LLMs, Claude 4.6, DeepSeek V4 Pro, Gemini 3.1 Pro, and GPT-5.5, were evaluated using a staged, rule-guided prompting workflow covering title screening, abstract screening, and full-text assessment. Model performance was assessed using accuracy, precision, recall, specificity, the F1-score, and Cohen's kappa, supplemented by stage-wise screening flow analysis and qualitative discrepancy analysis between model and human reviewer decisions. Findings Model performance varied substantially across datasets and models. In the conceptually diffuse Dataset A, all models achieved high specificity, but positive-class performance was more limited. DeepSeek V4 Pro achieved the highest accuracy, precision, F1-score, and Cohen's kappa, whereas Claude 4.6 achieved the highest recall. In the more clearly bounded Dataset B, recall was higher and model behavior was more convergent, with GPT-5.5 showing the best overall balance across performance indicators. Stage-wise analysis showed that models differed in where they made inclusion and exclusion decisions, while discrepancy analysis indicated that errors were mainly related to conceptual scope confusion, context misidentification, and underestimation of analytical depth. These findings suggest that current LLMs can support literature screening, but their reliability depends strongly on the conceptual clarity of the review task and the structure of eligibility criteria. Originality/value This study extends the literature on LLM-assisted screening by focusing on conceptually abstract and interdisciplinary review tasks. This study introduces a staged, Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA)-aligned screening workflow and conducts an error analysis that explains how LLM screening can fail in semantically diffuse settings. The study provides practical implications for the transparent and responsible use of LLMs in informetric research and systematic review practice.

Aslib Journal of Information Management
Huazhong Agricultural University (CN), EA Technology (GB), Hubei Provincial Center for Disease Control and Prevention (CN), Zhejiang University (CN)
Reduced inequalities
Openalex Percentile: Top 9%
Meta-analysis and systematic reviews
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.