Large language models for population-level public health communication: an evidence map of deployment, reach and policy gaps

Preprint version 7 (posted 16 September 2026). This version supersedes versions 1–6 and should be used in preference to them. What changed, and why. Version 7 reports a holdout-validated re-analysis. The automated charting instrument was validated against a disjoint 75-study human reference standard, charted by hand and blind to all machine output and frozen before scoring. Every accuracy figure in the manuscript is now the held-out estimate (for example, deployment stage 89.9% and target language 88.9% corpus-weighted exact match) rather than the earlier instrument-selection-sample figure, and each came in below its selection-sample counterpart, as expected for a held-out estimate. A model-family sensitivity analysis (encoder-only versus generative) was added; category-specific precision and recall replace the aggregate accuracy for the detection-to-correction comparison (the ratio widens from about 11:1 to about 15:1 under correction); screening agreement is reported on both the full validation sample (71.8%) and the definitive-decision subset (97.5%, Cohen’s κ=0.83); the year-end publication projection was removed; and the reference list was re-verified in full. All 552-study corpus counts are unchanged. Data and code. The charting datasets (uncorrected and corrected), screening decisions, both human reference standards (the 50-study selection sample and the 75-study holdout), the machine second-coder output, the accuracy computation and all analysis code are openly available on the Open Science Framework (https://doi.org/10.17605/OSF.IO/X5N78; project node https://osf.io/n8q4y/). Abstract. Large language models (LLMs) are increasingly proposed for population-level public health communication, yet whether this research is translating into deployed tools that reach the populations and policy settings that most need them is unknown. We conducted a computationally assisted evidence map (reported using the PRISMA extension for scoping reviews), screening 30 715 records from six databases and grey literature (2019–2026). AI-assisted screening agreed with a second screen on 71.8% of the 933 sampled records (counting uncertain flags as non-agreement) and 97.5% of definitive decisions (κ=0.83). Charting used a hybrid instrument selected against a 50-study reference standard and validated on a separate hand-charted 75-study holdout, giving 51% to 90% exact-match accuracy by variable; full text was available for 171 of the 552 included studies (31.0%), the rest charted from title and abstract. The field grew from 3 studies in 2020 to 215 in 2025, but has not translated into practice: 84.1% of studies remained at concept or laboratory stage, 8.5% reported user testing and 1.1% real-world deployment. It is also skewed away from global need. Among 474 studies with identifiable affiliations, first authorship concentrated in the United States (38%) and other high-income countries; 72.1% stated no target language in the available text, a rate the holdout confirms. Misinformation detection (235 studies) exceeded correction (21) by an order of magnitude, a ratio that widens under category-specific correction. Turning a technically productive field into deployed public-health capacity requires policy attention to implementation and deliberate investment in multilingual, LMIC-relevant and correction-oriented applications.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-16
DOI
https://doi.org/10.5281/zenodo.22782102
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Large language models for population-level public health communication: an evidence map of deployment, reach and policy gaps

Hayden Farquhar
Zenodo (CERN European Organization for Nuclear Research)
Artificial Intelligence in Healthcare and Education
preprint

Large language models for population-level public health communication: an evidence map of deployment, reach and policy gaps

Hayden Farquhar
preprint en

Abstract

Preprint version 7 (posted 16 September 2026). This version supersedes versions 1–6 and should be used in preference to them. What changed, and why. Version 7 reports a holdout-validated re-analysis. The automated charting instrument was validated against a disjoint 75-study human reference standard, charted by hand and blind to all machine output and frozen before scoring. Every accuracy figure in the manuscript is now the held-out estimate (for example, deployment stage 89.9% and target language 88.9% corpus-weighted exact match) rather than the earlier instrument-selection-sample figure, and each came in below its selection-sample counterpart, as expected for a held-out estimate. A model-family sensitivity analysis (encoder-only versus generative) was added; category-specific precision and recall replace the aggregate accuracy for the detection-to-correction comparison (the ratio widens from about 11:1 to about 15:1 under correction); screening agreement is reported on both the full validation sample (71.8%) and the definitive-decision subset (97.5%, Cohen’s κ=0.83); the year-end publication projection was removed; and the reference list was re-verified in full. All 552-study corpus counts are unchanged. Data and code. The charting datasets (uncorrected and corrected), screening decisions, both human reference standards (the 50-study selection sample and the 75-study holdout), the machine second-coder output, the accuracy computation and all analysis code are openly available on the Open Science Framework (https://doi.org/10.17605/OSF.IO/X5N78; project node https://osf.io/n8q4y/). Abstract. Large language models (LLMs) are increasingly proposed for population-level public health communication, yet whether this research is translating into deployed tools that reach the populations and policy settings that most need them is unknown. We conducted a computationally assisted evidence map (reported using the PRISMA extension for scoping reviews), screening 30 715 records from six databases and grey literature (2019–2026). AI-assisted screening agreed with a second screen on 71.8% of the 933 sampled records (counting uncertain flags as non-agreement) and 97.5% of definitive decisions (κ=0.83). Charting used a hybrid instrument selected against a 50-study reference standard and validated on a separate hand-charted 75-study holdout, giving 51% to 90% exact-match accuracy by variable; full text was available for 171 of the 552 included studies (31.0%), the rest charted from title and abstract. The field grew from 3 studies in 2020 to 215 in 2025, but has not translated into practice: 84.1% of studies remained at concept or laboratory stage, 8.5% reported user testing and 1.1% real-world deployment. It is also skewed away from global need. Among 474 studies with identifiable affiliations, first authorship concentrated in the United States (38%) and other high-income countries; 72.1% stated no target language in the available text, a rate the holdout confirms. Misinformation detection (235 studies) exceeded correction (21) by an order of magnitude, a ratio that widens under category-specific correction. Turning a technically productive field into deployed public-health capacity requires policy attention to implementation and deliberate investment in multilingual, LMIC-relevant and correction-oriented applications.

Zenodo (CERN European Organization for Nuclear Research)
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.