Transformer-based language models for population-level public health communication: an evidence map of deployment, reach and policy gaps

Evidence map of transformer-based language models applied to population-level public health communication: 552 studies from six databases and grey literature (2019–2026), reported using PRISMA-ScR against a prospectively registered protocol. Version 9 (24 September 2026). The upper limit on deployments missed by the charting instrument is recalculated according to the audit's sampling design: records examined in full contribute their observed count, and only the stratified random sample is used for inference to records not charted. Four full-text-confirmed deployments (0.7%) are the minimum and the upper limit is 15 of 552 (2.7%), replacing the pooled figure of nine (1.6%) given in version 8, which treated three differently designed samples as one random sample. Headline statements now refer to operational deployment documented in the included studies; the target-language finding is kept distinct from target population; a strict-rule model-family sensitivity analysis is added (Table S20); two citation placements are corrected and DOIs added to the reference list. Transformer-based language models, spanning encoder-only BERT-style architectures and generative large language models (LLMs), are increasingly proposed for population-level public health communication, yet whether they reach deployment or the populations that most need them is unknown. We conducted a computationally assisted evidence map (PRISMA-ScR) of 30 715 records from six databases and grey literature (2019–2026). Encoder-only and generative studies are reported separately where the distinction matters. AI-assisted screening agreed with a second screen on 71.8% of 933 sampled records (uncertain flags as non-agreement) and 97.5% of definitive decisions (κ=0.83). Charting used a hybrid instrument validated on a hand-charted 75-study holdout, giving 51–90% exact-match accuracy by variable; full text was available for 171 of 552 studies (31.0%), the rest charted from abstracts. The field grew from 3 studies in 2020 to 215 in 2025, but included studies rarely documented operational deployment: 84.1% remained at concept or laboratory stage, 8.5% reported user testing and four (0.7%; audit-based upper limit 2.7%) documented deployment on full-text verification. Two patterns bear on reach: first authorship concentrated in the United States (38% of 474 with identifiable affiliations) and other high-income countries, and 72.1% stated no target language in the available text, confirmed by the holdout. Misinformation detection (235) exceeded correction (21) by an order of magnitude, widening under category-specific correction. Turning a technically productive field into deployed public-health capacity requires policy attention to implementation and investment in multilingual, LMIC-relevant and correction-oriented work. This deposit contains the preprint, its supplementary materials (Tables S1–S23) and the six main figures as standalone files. The full data, charting instruments, validation apparatus, deployment-audit worksheets and the bound computation are on the linked Open Science Framework project.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-24
DOI
https://doi.org/10.5281/zenodo.22928583
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Transformer-based language models for population-level public health communication: an evidence map of deployment, reach and policy gaps

Hayden Farquhar
Zenodo (CERN European Organization for Nuclear Research)
Artificial Intelligence in Healthcare and Education
preprint

Transformer-based language models for population-level public health communication: an evidence map of deployment, reach and policy gaps

Hayden Farquhar
preprint en

Abstract

Evidence map of transformer-based language models applied to population-level public health communication: 552 studies from six databases and grey literature (2019–2026), reported using PRISMA-ScR against a prospectively registered protocol. Version 9 (24 September 2026). The upper limit on deployments missed by the charting instrument is recalculated according to the audit's sampling design: records examined in full contribute their observed count, and only the stratified random sample is used for inference to records not charted. Four full-text-confirmed deployments (0.7%) are the minimum and the upper limit is 15 of 552 (2.7%), replacing the pooled figure of nine (1.6%) given in version 8, which treated three differently designed samples as one random sample. Headline statements now refer to operational deployment documented in the included studies; the target-language finding is kept distinct from target population; a strict-rule model-family sensitivity analysis is added (Table S20); two citation placements are corrected and DOIs added to the reference list. Transformer-based language models, spanning encoder-only BERT-style architectures and generative large language models (LLMs), are increasingly proposed for population-level public health communication, yet whether they reach deployment or the populations that most need them is unknown. We conducted a computationally assisted evidence map (PRISMA-ScR) of 30 715 records from six databases and grey literature (2019–2026). Encoder-only and generative studies are reported separately where the distinction matters. AI-assisted screening agreed with a second screen on 71.8% of 933 sampled records (uncertain flags as non-agreement) and 97.5% of definitive decisions (κ=0.83). Charting used a hybrid instrument validated on a hand-charted 75-study holdout, giving 51–90% exact-match accuracy by variable; full text was available for 171 of 552 studies (31.0%), the rest charted from abstracts. The field grew from 3 studies in 2020 to 215 in 2025, but included studies rarely documented operational deployment: 84.1% remained at concept or laboratory stage, 8.5% reported user testing and four (0.7%; audit-based upper limit 2.7%) documented deployment on full-text verification. Two patterns bear on reach: first authorship concentrated in the United States (38% of 474 with identifiable affiliations) and other high-income countries, and 72.1% stated no target language in the available text, confirmed by the holdout. Misinformation detection (235) exceeded correction (21) by an order of magnitude, widening under category-specific correction. Turning a technically productive field into deployed public-health capacity requires policy attention to implementation and investment in multilingual, LMIC-relevant and correction-oriented work. This deposit contains the preprint, its supplementary materials (Tables S1–S23) and the six main figures as standalone files. The full data, charting instruments, validation apparatus, deployment-audit worksheets and the bound computation are on the linked Open Science Framework project.

Zenodo (CERN European Organization for Nuclear Research)
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.