Evaluating a Frozen LLM Configuration for Clinical-Guideline Literature Screening: A Protocol-Defined Validation Study

This preprint examines whether a fixed large language model configuration could safely support title-and-abstract screening in a planned evidence-synthesis workstream on clinical-guideline recommendations. The evaluation used a 1,048-record stratified calibration frame enriched for relevant records, rather than a representative sample of the larger search corpus. A dual-reviewer human reference standard identified 498 relevant records. Before model execution, the study protocol specified that the exact one-sided 95% lower confidence bound for recall had to reach 95% before the model could be used to make autonomous exclusion decisions. The configuration preserved 234 of the 498 relevant records for human review and routed 264 relevant records to exclusion. Observed recall was 46.99% (234/498), with a one-sided 95% Clopper-Pearson lower bound of 43.23%, below the required threshold. Therefore, the configuration was not considered safe for autonomous exclusion in this evaluation. A descriptive review of missed records suggested that many were related to clinical-guideline updating, implementation, comparison, or versioning, areas broader than the model's operational inclusion construct. These results apply only to this configuration, eligibility rubric, and calibration frame. They do not estimate performance on the full search corpus or support general conclusions about large language models in literature screening.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-18
DOI
https://doi.org/10.5281/zenodo.22838608
Primary Topic
Meta-analysis and systematic reviews
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Evaluating a Frozen LLM Configuration for Clinical-Guideline Literature Screening: A Protocol-Defined Validation Study

Mohamed Faisal Sindhi
Zenodo (CERN European Organization for Nuclear Research)
Meta-analysis and systematic reviews
preprint

Evaluating a Frozen LLM Configuration for Clinical-Guideline Literature Screening: A Protocol-Defined Validation Study

Mohamed Faisal Sindhi
preprint en

Abstract

This preprint examines whether a fixed large language model configuration could safely support title-and-abstract screening in a planned evidence-synthesis workstream on clinical-guideline recommendations. The evaluation used a 1,048-record stratified calibration frame enriched for relevant records, rather than a representative sample of the larger search corpus. A dual-reviewer human reference standard identified 498 relevant records. Before model execution, the study protocol specified that the exact one-sided 95% lower confidence bound for recall had to reach 95% before the model could be used to make autonomous exclusion decisions. The configuration preserved 234 of the 498 relevant records for human review and routed 264 relevant records to exclusion. Observed recall was 46.99% (234/498), with a one-sided 95% Clopper-Pearson lower bound of 43.23%, below the required threshold. Therefore, the configuration was not considered safe for autonomous exclusion in this evaluation. A descriptive review of missed records suggested that many were related to clinical-guideline updating, implementation, comparison, or versioning, areas broader than the model's operational inclusion construct. These results apply only to this configuration, eligibility rubric, and calibration frame. They do not estimate performance on the full search corpus or support general conclusions about large language models in literature screening.

Zenodo (CERN European Organization for Nuclear Research)
University of South Florida (US)
Reduced inequalities
Meta-analysis and systematic reviews
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Evaluating a Frozen LLM Configuration for Clinical-Guideline Literature Screening: A Protocol-Defined Validation Study — Mohamed Faisal Sindhi · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS