Protein-based resistance in-silico model (PRISM-TB) for rapid prediction of drug-resistant tuberculosis using machine learning
Background and objectives Tuberculosis (TB) caused by Mycobacterium tuberculosis (MTB), remains a priority health challenge with multidrug-resistant (MDR) strains threatening control efforts. Conventional diagnostic methods are limited by high cost, prolonged diagnostic time and incomplete mutation coverage. This study presents PRISM-TB (Protein-based resistance in-silico model for TB), a machine learning-based framework for predicting resistance to first- and second-line anti-TB drugs using protein sequence-derived features. Methods Protein sequences of 11 resistance-associated Mycobacterium tuberculosis genes were retrieved from NCBI and curated using literature-confirmed resistance mutations to generate a labelled dataset of 1,369 sequences across six anti-TB drugs. Sequence-derived numerical features were extracted and subjected to standardised preprocessing prior to model development. Six supervised machine-learning algorithms were trained and optimised using Bayesian hyperparameter tuning with stratified cross-validation, followed by model performance evaluation. External validation was performed using 461 mutations from the WHO Catalogue of Mutations in M. tuberculosis complex. Results Ensemble classifiers consistently outperformed linear and probabilistic models in predicting drug resistance from protein sequence derived features. Extra Trees achieved the highest overall performance (F1-score = 0.942, accuracy = 0.941), while Random Forest demonstrated superior class discrimination (ROC-AUC = 0.973) using only 10 features. On external validation, Random Forest correctly identified 351 of 374 WHO-graded resistant mutations (recall = 0.939, accuracy = 0.818). Feature importance and SHapely Additive exPlanations (SHAP) analyses identified physicochemical descriptors related to charge, hydrophobicity, polarity, and sequence transitions as the primary drivers of resistance prediction. Interpretation and conclusions This study highlights the potential of integrating protein-level mutation data with interpretable machine learning models for rapid, in-silico prediction of MDR-TB, offering a cost-effective approach to resistance screening.
Authors
- Achint Chaudhary
- SukhDev Mishra (ORCID: https://orcid.org/0000-0002-0842-1657)
- Uma Devi Ranganathan
- Swati Joshi
- Theja K.V.
- Ram Shankar Barai
- Agniva Das
Institutions
- Gujarat University (IN)
- Symbiosis International University (IN)
- Indian Council of Medical Research (IN)
- National Institute of Occupational Health (IN)
- National Institute of Research in Tuberculosis (IN)
- Academy of Scientific and Innovative Research (IN)
Publication Details
- Journal
- The Indian Journal of Medical Research
- Published
- 2026-09-30
- DOI
- https://doi.org/10.25259/ijmr_3532_2025
- Primary Topic
- Tuberculosis Research and Epidemiology
- Type
- article
- Field-Weighted Citation Impact
- 0.00