Using Natural Language Processing to Examine State Child Maltreatment Policies and Associations With Outcomes: Multistate Cross-Sectional Study

Background Child maltreatment is a major public health issue in the United States, with substantial variation in how states define, report, and respond to abuse and neglect. While prior research has examined individual policy components, less is known about how multiple policies jointly shape broader policy environments and relate to maltreatment outcomes. Objective This study used natural language processing (NLP) to characterize state child maltreatment policy environments and examine their associations with maltreatment incidence, recurrence, and fatalities across US jurisdictions. Methods A cross-sectional study was conducted using 2021 data from 50 US states, the District of Columbia, and Puerto Rico (N=52). A total of 411 state maltreatment policy items were derived from the State Child Abuse & Neglect Policies Database. A total of 6 NLP models (bidirectional and auto-regressive transformer [BART], bidirectional encoder representations from transformers [BERT], robustly optimized BERT approach [RoBERTa], decoding-enhanced BERT with disentangled attention [DeBERTa], Copilot, and LLaMA 3.1) were applied using a zero-shot classification framework to quantify policy characteristics. Models were evaluated using intrinsic (category consistency and semantic alignment) and extrinsic (factor analysis and clustering performance) metrics. Exploratory factor analysis was used to identify latent policy domains, and k-means clustering was applied to group jurisdictions with similar policy profiles. Maltreatment outcomes, including incidence, recurrence, and fatalities, were obtained from the National Child Abuse and Neglect Data System. Outcome differences across clusters were assessed using ANOVA with post hoc pairwise comparisons. Results The study population included 72,838,819 children and 3,774,528 maltreatment reports, of which 751,283 were substantiated or indicated cases. Across 6 NLP models, DeBERTa demonstrated the best overall performance. Exploratory factor analysis identified 3 primary policy domains: maltreatment definition, mandated reporting, and alternative response. A total of 4 policy clusters were identified. Jurisdictions with weaker reporting requirements and fewer penalties (cluster 2) had the highest incidence of maltreatment (mean 17.51, SD 6.13 vs mean 9.65, SD 4.10; mean 10.28, SD 6.47; and mean 11.74, SD 6.47 per 1000 children; P=.07) and recurrence (mean 1.74, SD 0.87 vs mean 0.52, SD 0.44; mean 0.66, SD 0.52; and mean 0.68, SD 0.43 per 1000 children; P<.001). Jurisdictions with stronger reporting requirements, broader definitions, and greater use of alternative response systems (cluster 1) had the lowest incidence and recurrence. No statistically significant differences in maltreatment fatalities were observed across clusters (P=.89). However, in a subanalysis of 18 fatality-related policy items, clusters differed significantly in fatality rates, with jurisdictions characterized by less clearly defined fatality policies exhibiting higher fatality rates (mean 47.91, SD 8.35 vs mean 22.64, SD 11.82; mean 25.63, SD 18.62; and mean 19.03, SD 11.64 per 1,000,000 children; P=.03). Conclusions State child maltreatment policies form distinct, multidimensional patterns associated with maltreatment outcomes. The findings highlight the importance of evaluating integrated policy frameworks rather than isolated components and demonstrate the use of NLP for large-scale, data-driven policy analysis.

Authors

Publication Details

Journal
JMIR Public Health and Surveillance
Published
2026-09-16
DOI
https://doi.org/10.2196/99643
Primary Topic
Child Abuse and Trauma
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Using Natural Language Processing to Examine State Child Maltreatment Policies and Associations With Outcomes: Multistate Cross-Sectional Study

Lutfiyya N. Muhammad, Richard A. Epstein, Neil Jordan, Nethra Sambamoorthi et al.
JMIR Public Health and Surveillance
Child Abuse and Trauma
article

Using Natural Language Processing to Examine State Child Maltreatment Policies and Associations With Outcomes: Multistate Cross-Sectional Study

Lutfiyya N. Muhammad, Richard A. Epstein, Neil Jordan, Nethra Sambamoorthi, Zhidi Luo
article en

Abstract

Background Child maltreatment is a major public health issue in the United States, with substantial variation in how states define, report, and respond to abuse and neglect. While prior research has examined individual policy components, less is known about how multiple policies jointly shape broader policy environments and relate to maltreatment outcomes. Objective This study used natural language processing (NLP) to characterize state child maltreatment policy environments and examine their associations with maltreatment incidence, recurrence, and fatalities across US jurisdictions. Methods A cross-sectional study was conducted using 2021 data from 50 US states, the District of Columbia, and Puerto Rico (N=52). A total of 411 state maltreatment policy items were derived from the State Child Abuse & Neglect Policies Database. A total of 6 NLP models (bidirectional and auto-regressive transformer [BART], bidirectional encoder representations from transformers [BERT], robustly optimized BERT approach [RoBERTa], decoding-enhanced BERT with disentangled attention [DeBERTa], Copilot, and LLaMA 3.1) were applied using a zero-shot classification framework to quantify policy characteristics. Models were evaluated using intrinsic (category consistency and semantic alignment) and extrinsic (factor analysis and clustering performance) metrics. Exploratory factor analysis was used to identify latent policy domains, and k-means clustering was applied to group jurisdictions with similar policy profiles. Maltreatment outcomes, including incidence, recurrence, and fatalities, were obtained from the National Child Abuse and Neglect Data System. Outcome differences across clusters were assessed using ANOVA with post hoc pairwise comparisons. Results The study population included 72,838,819 children and 3,774,528 maltreatment reports, of which 751,283 were substantiated or indicated cases. Across 6 NLP models, DeBERTa demonstrated the best overall performance. Exploratory factor analysis identified 3 primary policy domains: maltreatment definition, mandated reporting, and alternative response. A total of 4 policy clusters were identified. Jurisdictions with weaker reporting requirements and fewer penalties (cluster 2) had the highest incidence of maltreatment (mean 17.51, SD 6.13 vs mean 9.65, SD 4.10; mean 10.28, SD 6.47; and mean 11.74, SD 6.47 per 1000 children; P=.07) and recurrence (mean 1.74, SD 0.87 vs mean 0.52, SD 0.44; mean 0.66, SD 0.52; and mean 0.68, SD 0.43 per 1000 children; P<.001). Jurisdictions with stronger reporting requirements, broader definitions, and greater use of alternative response systems (cluster 1) had the lowest incidence and recurrence. No statistically significant differences in maltreatment fatalities were observed across clusters (P=.89). However, in a subanalysis of 18 fatality-related policy items, clusters differed significantly in fatality rates, with jurisdictions characterized by less clearly defined fatality policies exhibiting higher fatality rates (mean 47.91, SD 8.35 vs mean 22.64, SD 11.82; mean 25.63, SD 18.62; and mean 19.03, SD 11.64 per 1,000,000 children; P=.03). Conclusions State child maltreatment policies form distinct, multidimensional patterns associated with maltreatment outcomes. The findings highlight the importance of evaluating integrated policy frameworks rather than isolated components and demonstrate the use of NLP for large-scale, data-driven policy analysis.

JMIR Public Health and SurveillanceVol. 12
Peace, Justice and strong institutions
Openalex Percentile: Top 7%
Child Abuse and Trauma
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.