Workflow-induced evaluation targets in automation-assisted screening: Implications for cumulative evidence in research synthesis

Abstract Automation-assisted title–abstract screening is routinely evaluated using metrics such as recall, precision, and workload reduction, and results are increasingly summarized across studies. Synthesis presupposes that evaluation results refer to comparable target quantities. This article examines whether screening evaluation results produced under diverse automation-assisted workflows can meaningfully support cumulative inference, and under what conditions cross-study comparison is warranted. This conceptual article treats screening evaluation targets as workflow-induced rather than task-inherent. Evaluation design is decomposed into four workflow dimensions ( representation , inference , governance , and evaluation ) plus the reference-decision set produced or specified. Recurring choice patterns across these elements define evaluation configurations. Two study-level configurations and one synthesis-level pattern are illustrated: endogenous adjudication , inherited adjudication , and aggregation of heterogeneous targets . Across configurations, evaluation targets are induced by workflow-specific reference-decision sets and, in some designs, model-based target domains rather than by a common reference set. Nominally similar metrics can therefore estimate different quantities across studies, and apparent cumulativeness can arise from pooling estimates of different quantities. Standard heterogeneity statistics cannot distinguish variation in method effectiveness from variation in what is being estimated. Comparison is warranted only when targets align before pooling. Differences across screening evaluations may reflect differences in what is being estimated rather than statistical heterogeneity around a common estimand. When evaluation targets are not aligned across the framework, aggregate summaries can describe reported results but cannot support cumulative inference. Improved aggregation requires both better study-level specification of evaluation targets and explicit attention to configuration variation at synthesis .

Authors

Institutions

Publication Details

Journal
Research Synthesis Methods
Published
2026-09-21
DOI
https://doi.org/10.1017/rsm.2026.10116
Primary Topic
Scientific Computing and Data Management
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Workflow-induced evaluation targets in automation-assisted screening: Implications for cumulative evidence in research synthesis

John Alan Nunnery
Research Synthesis Methods
Scientific Computing and Data Management
article

Workflow-induced evaluation targets in automation-assisted screening: Implications for cumulative evidence in research synthesis

John Alan Nunnery
article en

Abstract

Abstract Automation-assisted title–abstract screening is routinely evaluated using metrics such as recall, precision, and workload reduction, and results are increasingly summarized across studies. Synthesis presupposes that evaluation results refer to comparable target quantities. This article examines whether screening evaluation results produced under diverse automation-assisted workflows can meaningfully support cumulative inference, and under what conditions cross-study comparison is warranted. This conceptual article treats screening evaluation targets as workflow-induced rather than task-inherent. Evaluation design is decomposed into four workflow dimensions ( representation , inference , governance , and evaluation ) plus the reference-decision set produced or specified. Recurring choice patterns across these elements define evaluation configurations. Two study-level configurations and one synthesis-level pattern are illustrated: endogenous adjudication , inherited adjudication , and aggregation of heterogeneous targets . Across configurations, evaluation targets are induced by workflow-specific reference-decision sets and, in some designs, model-based target domains rather than by a common reference set. Nominally similar metrics can therefore estimate different quantities across studies, and apparent cumulativeness can arise from pooling estimates of different quantities. Standard heterogeneity statistics cannot distinguish variation in method effectiveness from variation in what is being estimated. Comparison is warranted only when targets align before pooling. Differences across screening evaluations may reflect differences in what is being estimated rather than statistical heterogeneity around a common estimand. When evaluation targets are not aligned across the framework, aggregate summaries can describe reported results but cannot support cumulative inference. Improved aggregation requires both better study-level specification of evaluation targets and explicit attention to configuration variation at synthesis .

Research Synthesis Methods
Dominion University College (GH), Old Dominion University (US)
Openalex Percentile: Top 4%
Scientific Computing and Data Management
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.