Transformer-based language models for population-level public health communication: an evidence map of deployment, reach and policy gaps
Evidence map of transformer-based language models applied to population-level public health communication: 552 studies from six databases and grey literature (2019–2026), reported using PRISMA-ScR against a prospectively registered protocol. Version 9 (24 September 2026). The upper limit on deployments missed by the charting instrument is recalculated according to the audit's sampling design: records examined in full contribute their observed count, and only the stratified random sample is used for inference to records not charted. Four full-text-confirmed deployments (0.7%) are the minimum and the upper limit is 15 of 552 (2.7%), replacing the pooled figure of nine (1.6%) given in version 8, which treated three differently designed samples as one random sample. Headline statements now refer to operational deployment documented in the included studies; the target-language finding is kept distinct from target population; a strict-rule model-family sensitivity analysis is added (Table S20); two citation placements are corrected and DOIs added to the reference list. Transformer-based language models, spanning encoder-only BERT-style architectures and generative large language models (LLMs), are increasingly proposed for population-level public health communication, yet whether they reach deployment or the populations that most need them is unknown. We conducted a computationally assisted evidence map (PRISMA-ScR) of 30 715 records from six databases and grey literature (2019–2026). Encoder-only and generative studies are reported separately where the distinction matters. AI-assisted screening agreed with a second screen on 71.8% of 933 sampled records (uncertain flags as non-agreement) and 97.5% of definitive decisions (κ=0.83). Charting used a hybrid instrument validated on a hand-charted 75-study holdout, giving 51–90% exact-match accuracy by variable; full text was available for 171 of 552 studies (31.0%), the rest charted from abstracts. The field grew from 3 studies in 2020 to 215 in 2025, but included studies rarely documented operational deployment: 84.1% remained at concept or laboratory stage, 8.5% reported user testing and four (0.7%; audit-based upper limit 2.7%) documented deployment on full-text verification. Two patterns bear on reach: first authorship concentrated in the United States (38% of 474 with identifiable affiliations) and other high-income countries, and 72.1% stated no target language in the available text, confirmed by the holdout. Misinformation detection (235) exceeded correction (21) by an order of magnitude, widening under category-specific correction. Turning a technically productive field into deployed public-health capacity requires policy attention to implementation and investment in multilingual, LMIC-relevant and correction-oriented work. This deposit contains the preprint, its supplementary materials (Tables S1–S23) and the six main figures as standalone files. The full data, charting instruments, validation apparatus, deployment-audit worksheets and the bound computation are on the linked Open Science Framework project.
Authors
- Hayden Farquhar (ORCID: https://orcid.org/0009-0002-6226-440X)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-24
- DOI
- https://doi.org/10.5281/zenodo.22928583
- Primary Topic
- Artificial Intelligence in Healthcare and Education
- Type
- preprint