Large language models for population-level public health communication: an evidence map of deployment, reach and policy gaps
Preprint version 7 (posted 16 September 2026). This version supersedes versions 1–6 and should be used in preference to them. What changed, and why. Version 7 reports a holdout-validated re-analysis. The automated charting instrument was validated against a disjoint 75-study human reference standard, charted by hand and blind to all machine output and frozen before scoring. Every accuracy figure in the manuscript is now the held-out estimate (for example, deployment stage 89.9% and target language 88.9% corpus-weighted exact match) rather than the earlier instrument-selection-sample figure, and each came in below its selection-sample counterpart, as expected for a held-out estimate. A model-family sensitivity analysis (encoder-only versus generative) was added; category-specific precision and recall replace the aggregate accuracy for the detection-to-correction comparison (the ratio widens from about 11:1 to about 15:1 under correction); screening agreement is reported on both the full validation sample (71.8%) and the definitive-decision subset (97.5%, Cohen’s κ=0.83); the year-end publication projection was removed; and the reference list was re-verified in full. All 552-study corpus counts are unchanged. Data and code. The charting datasets (uncorrected and corrected), screening decisions, both human reference standards (the 50-study selection sample and the 75-study holdout), the machine second-coder output, the accuracy computation and all analysis code are openly available on the Open Science Framework (https://doi.org/10.17605/OSF.IO/X5N78; project node https://osf.io/n8q4y/). Abstract. Large language models (LLMs) are increasingly proposed for population-level public health communication, yet whether this research is translating into deployed tools that reach the populations and policy settings that most need them is unknown. We conducted a computationally assisted evidence map (reported using the PRISMA extension for scoping reviews), screening 30 715 records from six databases and grey literature (2019–2026). AI-assisted screening agreed with a second screen on 71.8% of the 933 sampled records (counting uncertain flags as non-agreement) and 97.5% of definitive decisions (κ=0.83). Charting used a hybrid instrument selected against a 50-study reference standard and validated on a separate hand-charted 75-study holdout, giving 51% to 90% exact-match accuracy by variable; full text was available for 171 of the 552 included studies (31.0%), the rest charted from title and abstract. The field grew from 3 studies in 2020 to 215 in 2025, but has not translated into practice: 84.1% of studies remained at concept or laboratory stage, 8.5% reported user testing and 1.1% real-world deployment. It is also skewed away from global need. Among 474 studies with identifiable affiliations, first authorship concentrated in the United States (38%) and other high-income countries; 72.1% stated no target language in the available text, a rate the holdout confirms. Misinformation detection (235 studies) exceeded correction (21) by an order of magnitude, a ratio that widens under category-specific correction. Turning a technically productive field into deployed public-health capacity requires policy attention to implementation and deliberate investment in multilingual, LMIC-relevant and correction-oriented applications.
Authors
- Hayden Farquhar (ORCID: https://orcid.org/0009-0002-6226-440X)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-16
- DOI
- https://doi.org/10.5281/zenodo.22782102
- Primary Topic
- Artificial Intelligence in Healthcare and Education
- Type
- preprint