Phenotypic features outperform raw geolocation for community level cognitive difficulty prediction using BRFSS SMART data

Abstract Public health surveillance datasets provide population level signals that may complement clinical data for understanding dementia risk and allocating resources, offering scalable community level coverage that individual clinical records cannot match on their own. In this study, we investigate whether geolocation features (latitude/longitude of metropolitan and micropolitan statistical areas) add predictive value beyond phenotypic (health behavior and condition prevalence) signals when modeling a dementia relevant outcome at the community level. Using CDC’s BRFSS SMART MMSA prevalence dataset, we construct a modeling table (4,289 subgroup location observations across 190 MMSAs in 2023) with three disability status targets, including serious difficulty concentrating, remembering, or making decisions. We formulate a question conditioned regression problem: given the textual description of the target question, phenotypic prevalence features, and optionally geolocation, we predict the observed prevalence (). We evaluate TF-IDF and Sentence-BERT encoders combined with linear models, elastic net, random forests, XGBoost, and LightGBM, using fold wise (train only) imputation to avoid cross fold leakage. Under random 10-fold cross validation, adding geolocation improves performance over text only (R $$^2$$ : 0.598 $$\\rightarrow$$ 0.862), but the additional gain beyond text+phenotype is not statistically significant (0.915 $$\\rightarrow$$ 0.916; paired Wilcoxon $$p=0.70$$ ). Under spatial cross validation, used to test generalization to entirely unseen geographic regions, text+phenotype and text+geo+phenotype remain statistically indistinguishable (0.873 vs. 0.873; $$p=0.63$$ ), a result that is robust to the choice of spatial fold definition. A repeated permutation test (30 shuffles of latitude/longitude) likewise found no significant drop in performance ( $$p=0.39$$ ), indicating no reliable evidence that the model uses geolocation beyond noise. SHAP analysis highlights mobility limitation and mental health prevalence as dominant phenotypic contributors; latitude enters the top features when geolocation is included but does not dominate.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-09-17
DOI
https://doi.org/10.1038/s41598-026-72006-y
Primary Topic
Dementia and Cognitive Impairment Research
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Phenotypic features outperform raw geolocation for community level cognitive difficulty prediction using BRFSS SMART data

Divya Chaudhary, Peng Zhang
Scientific Reports
Dementia and Cognitive Impairment Research
article

Phenotypic features outperform raw geolocation for community level cognitive difficulty prediction using BRFSS SMART data

Divya Chaudhary, Peng Zhang
article en

Abstract

Abstract Public health surveillance datasets provide population level signals that may complement clinical data for understanding dementia risk and allocating resources, offering scalable community level coverage that individual clinical records cannot match on their own. In this study, we investigate whether geolocation features (latitude/longitude of metropolitan and micropolitan statistical areas) add predictive value beyond phenotypic (health behavior and condition prevalence) signals when modeling a dementia relevant outcome at the community level. Using CDC’s BRFSS SMART MMSA prevalence dataset, we construct a modeling table (4,289 subgroup location observations across 190 MMSAs in 2023) with three disability status targets, including serious difficulty concentrating, remembering, or making decisions. We formulate a question conditioned regression problem: given the textual description of the target question, phenotypic prevalence features, and optionally geolocation, we predict the observed prevalence (). We evaluate TF-IDF and Sentence-BERT encoders combined with linear models, elastic net, random forests, XGBoost, and LightGBM, using fold wise (train only) imputation to avoid cross fold leakage. Under random 10-fold cross validation, adding geolocation improves performance over text only (R $$^2$$ : 0.598 $$\rightarrow$$ 0.862), but the additional gain beyond text+phenotype is not statistically significant (0.915 $$\rightarrow$$ 0.916; paired Wilcoxon $$p=0.70$$ ). Under spatial cross validation, used to test generalization to entirely unseen geographic regions, text+phenotype and text+geo+phenotype remain statistically indistinguishable (0.873 vs. 0.873; $$p=0.63$$ ), a result that is robust to the choice of spatial fold definition. A repeated permutation test (30 shuffles of latitude/longitude) likewise found no significant drop in performance ( $$p=0.39$$ ), indicating no reliable evidence that the model uses geolocation beyond noise. SHAP analysis highlights mobility limitation and mental health prevalence as dominant phenotypic contributors; latitude enters the top features when geolocation is included but does not dominate.

Scientific Reports
Seattle University (US)
Openalex Percentile: Top 10%
Dementia and Cognitive Impairment Research
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.