Auditing Geographic Domain Shift in Genomic AI: A Memory-Efficient Pipeline for Tuberculosis Drug Resistance Prediction

The deployment of machine learning models in clinical genomics holds immense potential for rapid drug susceptibility testing (DST). However, clinical AI is highly vulnerable to geographic domain shift. This paper audits the geographic generalizability of XGBoost models predicting Isoniazid (INH) resistance in Mycobacterium tuberculosis. Utilizing the CRyPTIC dataset, we propose a memory-constrained preprocessing pipeline to extract and reshape high-volume genomic data, evaluate performance across distinct global laboratories, and deploy SHAP (SHapley Additive exPlanations) to ensure predictions are driven by biological mechanisms rather than geographic metadata artifacts. Our evaluation reveals a reverse domain shift, demonstrating that models can achieve superior performance on unseen populations with lower intrinsic genetic variance.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-05
DOI
https://doi.org/10.5281/zenodo.23162573
Primary Topic
Tuberculosis Research and Epidemiology
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Auditing Geographic Domain Shift in Genomic AI: A Memory-Efficient Pipeline for Tuberculosis Drug Resistance Prediction

Mohammad Kawsar
Zenodo (CERN European Organization for Nuclear Research)
Tuberculosis Research and Epidemiology
preprint

Auditing Geographic Domain Shift in Genomic AI: A Memory-Efficient Pipeline for Tuberculosis Drug Resistance Prediction

Mohammad Kawsar
preprint en

Abstract

The deployment of machine learning models in clinical genomics holds immense potential for rapid drug susceptibility testing (DST). However, clinical AI is highly vulnerable to geographic domain shift. This paper audits the geographic generalizability of XGBoost models predicting Isoniazid (INH) resistance in Mycobacterium tuberculosis. Utilizing the CRyPTIC dataset, we propose a memory-constrained preprocessing pipeline to extract and reshape high-volume genomic data, evaluate performance across distinct global laboratories, and deploy SHAP (SHapley Additive exPlanations) to ensure predictions are driven by biological mechanisms rather than geographic metadata artifacts. Our evaluation reveals a reverse domain shift, demonstrating that models can achieve superior performance on unseen populations with lower intrinsic genetic variance.

Zenodo (CERN European Organization for Nuclear Research)
Tuberculosis Research and Epidemiology
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.