Evaluating a Frozen LLM Configuration for Clinical-Guideline Literature Screening: A Protocol-Defined Validation Study
This preprint examines whether a fixed large language model configuration could safely support title-and-abstract screening in a planned evidence-synthesis workstream on clinical-guideline recommendations. The evaluation used a 1,048-record stratified calibration frame enriched for relevant records, rather than a representative sample of the larger search corpus. A dual-reviewer human reference standard identified 498 relevant records. Before model execution, the study protocol specified that the exact one-sided 95% lower confidence bound for recall had to reach 95% before the model could be used to make autonomous exclusion decisions. The configuration preserved 234 of the 498 relevant records for human review and routed 264 relevant records to exclusion. Observed recall was 46.99% (234/498), with a one-sided 95% Clopper-Pearson lower bound of 43.23%, below the required threshold. Therefore, the configuration was not considered safe for autonomous exclusion in this evaluation. A descriptive review of missed records suggested that many were related to clinical-guideline updating, implementation, comparison, or versioning, areas broader than the model's operational inclusion construct. These results apply only to this configuration, eligibility rubric, and calibration frame. They do not estimate performance on the full search corpus or support general conclusions about large language models in literature screening.
Authors
- Mohamed Faisal Sindhi
Institutions
- University of South Florida (US)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-18
- DOI
- https://doi.org/10.5281/zenodo.22838608
- Primary Topic
- Meta-analysis and systematic reviews
- Type
- preprint