Household-level food-insecurity prediction from experience-scale surveys and open geospatial data: A reproducible machine-learning pipeline for Nigeria
This paper develops a reproducible machine-learning pipeline for household-level food-insecurity prediction in Nigeria using the 2024 NDHS FIES module and five open geospatial layers. Because the NDHS is a two-stage cluster sample, all evaluation uses cluster-grouped data splits in which no enumeration area contributes households to both training and test sets. On a held-out test set of 7,999 households from 276 enumeration areas, the random forest reaches AUROC 0.727 (cluster-bootstrap 95 percent CI 0.704 to 0.747) and gradient-boosted trees 0.714 (0.690 to 0.737), against 0.679 (0.653 to 0.702) for a logistic baseline. Under leave-zone-out validation, which scores entire geopolitical zones the model never saw, the logistic baseline outperforms both ensembles in all six zones. SHAP attribution assigns between 42 and 52 percent of predictive signal to contextual predictors depending on model and coordinate treatment, a larger share than prior DHS tree-ensemble work reports. Travel time to the nearest city is associated with lower food-insecurity risk, a pattern consistent with a subsistence-buffer channel during sharp naira depreciation that we present as a hypothesis requiring replication.
Authors
- Philip Osung Osung (ORCID: https://orcid.org/0009-0001-6700-372X)
Institutions
- Nile University of Nigeria (NG)
Publication Details
- Journal
- PLoS ONE
- Published
- 2026-10-07
- DOI
- https://doi.org/10.1371/journal.pone.0349978
- Primary Topic
- Food Security and Socioeconomic Dynamics
- Type
- article
- Field-Weighted Citation Impact
- 0.00