Automating narrative analysis to enhance computer-coded verbal autopsies: a scoping review

Abstract Mortality surveillance, including ascertainment of fact and cause of death, is crucial for understanding disease trends and informing public health interventions and policies. Verbal autopsy (VA) interviews conducted by trained lay interviewers with the next of kin of the deceased are increasingly being used to determine causes of death (CoD), where a substantial proportion of deaths occur outside of clinical care settings. When reviewed and coded into ICD-based causes of death by physicians [Physician-Coded Verbal Autopsy (PCVA)] or by algorithms based on physician coding [Computer Coded Verbal Autopsies (CCVA)], VA data can complement mortality statistics coming from medically certified causes of death (MCCD). CCVA algorithms use discrete multiple-choice questions in the VA instrument to ascertain the probable cause of death. These algorithms have yet to fully utilize information in the narrative section of the VA. Processing unstructured and context-specific narrative data and integrating it into CCVA is challenging. In this scoping review, we review and synthesise existing literature on automation techniques applied to compute CoD using VA narratives, including Machine Learning models such as Bidirectional Encoder Representations from Transformers (BERT), Random Forest, A feedforward neural network (FNN), Extreme Gradient Boosting (XGBoost), and Support Vector Machine (SVM), and common Natural Language Processing (NLP) approaches including Term Frequency-Inverse Document Frequency (TF-IDF) and topic modeling methods such as Latent Dirichlet Allocation (LDA) that transform unstructured text into structured formats before CoD computation. Based on published studies, this review highlights the challenges and implications of using these methods. It underscores the need for open, up-to-date datasets to facilitate model comparison and improvement. Our findings suggest that integrating narratives into CoD prediction shows promise, particularly when combined with structured questionnaire data in hybrid modeling approaches. However, performance varied substantially across studies due to differences in datasets, preprocessing strategies, feature representation methods, and evaluation procedures. Although narrative-based models alone often underperform those developed from structured questions, incorporating narrative enhances CoD prediction by capturing context, chronology, and semantics, with some studies reporting improved classification performance compared to models using a single data modality.

Authors

Institutions

Publication Details

Journal
Population Health Metrics
Published
2026-09-21
DOI
https://doi.org/10.1186/s12963-026-00507-z
Primary Topic
Machine Learning in Healthcare
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Automating narrative analysis to enhance computer-coded verbal autopsies: a scoping review

Mahadia Tunga, Sangyal Dorjee, Isaac Lyatuu, Erin Nichols et al.
Population Health Metrics
Machine Learning in Healthcare
article

Automating narrative analysis to enhance computer-coded verbal autopsies: a scoping review

Mahadia Tunga, Sangyal Dorjee, Isaac Lyatuu, Erin Nichols, Melissa Marx, Daniel Cobos
article en

Abstract

Abstract Mortality surveillance, including ascertainment of fact and cause of death, is crucial for understanding disease trends and informing public health interventions and policies. Verbal autopsy (VA) interviews conducted by trained lay interviewers with the next of kin of the deceased are increasingly being used to determine causes of death (CoD), where a substantial proportion of deaths occur outside of clinical care settings. When reviewed and coded into ICD-based causes of death by physicians [Physician-Coded Verbal Autopsy (PCVA)] or by algorithms based on physician coding [Computer Coded Verbal Autopsies (CCVA)], VA data can complement mortality statistics coming from medically certified causes of death (MCCD). CCVA algorithms use discrete multiple-choice questions in the VA instrument to ascertain the probable cause of death. These algorithms have yet to fully utilize information in the narrative section of the VA. Processing unstructured and context-specific narrative data and integrating it into CCVA is challenging. In this scoping review, we review and synthesise existing literature on automation techniques applied to compute CoD using VA narratives, including Machine Learning models such as Bidirectional Encoder Representations from Transformers (BERT), Random Forest, A feedforward neural network (FNN), Extreme Gradient Boosting (XGBoost), and Support Vector Machine (SVM), and common Natural Language Processing (NLP) approaches including Term Frequency-Inverse Document Frequency (TF-IDF) and topic modeling methods such as Latent Dirichlet Allocation (LDA) that transform unstructured text into structured formats before CoD computation. Based on published studies, this review highlights the challenges and implications of using these methods. It underscores the need for open, up-to-date datasets to facilitate model comparison and improvement. Our findings suggest that integrating narratives into CoD prediction shows promise, particularly when combined with structured questionnaire data in hybrid modeling approaches. However, performance varied substantially across studies due to differences in datasets, preprocessing strategies, feature representation methods, and evaluation procedures. Although narrative-based models alone often underperform those developed from structured questions, incorporating narrative enhances CoD prediction by capturing context, chronology, and semantics, with some studies reporting improved classification performance compared to models using a single data modality.

Population Health Metrics
Ifakara Health Institute (TZ), National Center for Health Statistics (US), Johns Hopkins University (US), Swiss Tropical and Public Health Institute (CH), University of Dar es Salaam (TZ)
Good health and well-being
Openalex Percentile: Top 8%
Machine Learning in Healthcare
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.