MechBBB: A Two-Stage Mechanism Informed Machine Learning Tool for Blood Brain Barrier Permeability Prediction
Background/Objectives: Blood–brain barrier (BBB) permeability prediction is important for central nervous system drug discovery. Many models rely mainly on chemical structure, omit transport information, and may report optimistic performance when related compounds occur across training and test sets without adequate leakage control. We developed MechBBB, a two-stage, transport-informed machine-learning framework that adds learned efflux, influx, and passive-permeability scores to conventional molecular features and evaluates their added value under explicit leakage control. Methods: In Stage 1, three LightGBM models were trained on separate efflux, influx, and passive-permeability datasets to generate transport-related scores. BBBP compounds were excluded from Stage 1 training by InChIKey matching. In Stage 2, a LightGBM classifier was trained on BBBP using a Murcko scaffold split with the three Stage 1 scores, ten physicochemical descriptors, and 2048-bit ECFP4 fingerprints. Performance was compared with a descriptor-plus-fingerprint baseline without transport scores. External evaluation used a strict B3DB set (n = 4080) with no InChIKey or nonempty scaffold overlap with BBBP and was performed without retraining, recalibration, or threshold adjustment. MechBBB was also compared with official SwissADME BOILED-Egg predictions on a chemically matched external subset. Results: On the BBBP scaffold test set, MechBBB achieved an AUROC of 0.932, an AUPRC of 0.983, an MCC of 0.737, and a balanced accuracy of 0.866. The AUROC improvement over the descriptor-plus-fingerprint baseline was small but statistically significant in the paired test (ΔAUROC = 0.0096; p = 0.0425). Scaffold-grouped cross-validation did not show a consistent ranking advantage from adding the Stage 1 scores. On strict B3DB, MechBBB achieved an AUROC of 0.894 and an AUPRC of 0.893, with no significant ranking advantage over the baseline. On the matched B3DB subset (n = 4043), MechBBB achieved higher accuracy (0.814 vs. 0.682), balanced accuracy (0.806 vs. 0.693), sensitivity (0.910 vs. 0.541), and MCC (0.631 vs. 0.401) than SwissADME BOILED-Egg, whereas BOILED-Egg had higher specificity (0.845 vs. 0.703). SHAP analysis identified TPSA, NumHDonors, and p_pampa among the leading contributors. Conclusions: MechBBB provides a leakage controlled and externally tested framework that combines competitive BBB permeability prediction with per-compound transport-related information. The Stage 1 scores produced a modest internal improvement but did not provide a consistent external ranking advantage, indicating that their main added value is transport-related context rather than a universal increase in predictive accuracy. MechBBB can be used to prioritize compounds for CNS drug discovery before experimental BBB testing and identify compounds that may warrant follow-up studies of passive permeability or active transport. The MechBBB web tool is freely available.
Authors
- Sivanesan Dakshanamurthy (ORCID: https://orcid.org/0000-0003-3681-4121)
- Sahith Mada (ORCID: https://orcid.org/0009-0000-2606-2259)
- Yu Shin
Institutions
- Georgetown University (US)
- University of Michigan (US)
- Georgetown University Medical Center (US)
- University of Maryland, College Park (US)
Publication Details
- Journal
- Pharmaceuticals
- Published
- 2026-09-25
- DOI
- https://doi.org/10.3390/ph19101524
- Primary Topic
- Computational Drug Discovery Methods
- Type
- article
- Field-Weighted Citation Impact
- 0.00