MechBBB: A Two-Stage Mechanism Informed Machine Learning Tool for Blood Brain Barrier Permeability Prediction

Background/Objectives: Blood–brain barrier (BBB) permeability prediction is important for central nervous system drug discovery. Many models rely mainly on chemical structure, omit transport information, and may report optimistic performance when related compounds occur across training and test sets without adequate leakage control. We developed MechBBB, a two-stage, transport-informed machine-learning framework that adds learned efflux, influx, and passive-permeability scores to conventional molecular features and evaluates their added value under explicit leakage control. Methods: In Stage 1, three LightGBM models were trained on separate efflux, influx, and passive-permeability datasets to generate transport-related scores. BBBP compounds were excluded from Stage 1 training by InChIKey matching. In Stage 2, a LightGBM classifier was trained on BBBP using a Murcko scaffold split with the three Stage 1 scores, ten physicochemical descriptors, and 2048-bit ECFP4 fingerprints. Performance was compared with a descriptor-plus-fingerprint baseline without transport scores. External evaluation used a strict B3DB set (n = 4080) with no InChIKey or nonempty scaffold overlap with BBBP and was performed without retraining, recalibration, or threshold adjustment. MechBBB was also compared with official SwissADME BOILED-Egg predictions on a chemically matched external subset. Results: On the BBBP scaffold test set, MechBBB achieved an AUROC of 0.932, an AUPRC of 0.983, an MCC of 0.737, and a balanced accuracy of 0.866. The AUROC improvement over the descriptor-plus-fingerprint baseline was small but statistically significant in the paired test (ΔAUROC = 0.0096; p = 0.0425). Scaffold-grouped cross-validation did not show a consistent ranking advantage from adding the Stage 1 scores. On strict B3DB, MechBBB achieved an AUROC of 0.894 and an AUPRC of 0.893, with no significant ranking advantage over the baseline. On the matched B3DB subset (n = 4043), MechBBB achieved higher accuracy (0.814 vs. 0.682), balanced accuracy (0.806 vs. 0.693), sensitivity (0.910 vs. 0.541), and MCC (0.631 vs. 0.401) than SwissADME BOILED-Egg, whereas BOILED-Egg had higher specificity (0.845 vs. 0.703). SHAP analysis identified TPSA, NumHDonors, and p_pampa among the leading contributors. Conclusions: MechBBB provides a leakage controlled and externally tested framework that combines competitive BBB permeability prediction with per-compound transport-related information. The Stage 1 scores produced a modest internal improvement but did not provide a consistent external ranking advantage, indicating that their main added value is transport-related context rather than a universal increase in predictive accuracy. MechBBB can be used to prioritize compounds for CNS drug discovery before experimental BBB testing and identify compounds that may warrant follow-up studies of passive permeability or active transport. The MechBBB web tool is freely available.

Authors

Institutions

Publication Details

Journal
Pharmaceuticals
Published
2026-09-25
DOI
https://doi.org/10.3390/ph19101524
Primary Topic
Computational Drug Discovery Methods
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

MechBBB: A Two-Stage Mechanism Informed Machine Learning Tool for Blood Brain Barrier Permeability Prediction

Sivanesan Dakshanamurthy, Sahith Mada, Yu Shin
Pharmaceuticals
Computational Drug Discovery Methods
article

MechBBB: A Two-Stage Mechanism Informed Machine Learning Tool for Blood Brain Barrier Permeability Prediction

Sivanesan Dakshanamurthy, Sahith Mada, Yu Shin
article en

Abstract

Background/Objectives: Blood–brain barrier (BBB) permeability prediction is important for central nervous system drug discovery. Many models rely mainly on chemical structure, omit transport information, and may report optimistic performance when related compounds occur across training and test sets without adequate leakage control. We developed MechBBB, a two-stage, transport-informed machine-learning framework that adds learned efflux, influx, and passive-permeability scores to conventional molecular features and evaluates their added value under explicit leakage control. Methods: In Stage 1, three LightGBM models were trained on separate efflux, influx, and passive-permeability datasets to generate transport-related scores. BBBP compounds were excluded from Stage 1 training by InChIKey matching. In Stage 2, a LightGBM classifier was trained on BBBP using a Murcko scaffold split with the three Stage 1 scores, ten physicochemical descriptors, and 2048-bit ECFP4 fingerprints. Performance was compared with a descriptor-plus-fingerprint baseline without transport scores. External evaluation used a strict B3DB set (n = 4080) with no InChIKey or nonempty scaffold overlap with BBBP and was performed without retraining, recalibration, or threshold adjustment. MechBBB was also compared with official SwissADME BOILED-Egg predictions on a chemically matched external subset. Results: On the BBBP scaffold test set, MechBBB achieved an AUROC of 0.932, an AUPRC of 0.983, an MCC of 0.737, and a balanced accuracy of 0.866. The AUROC improvement over the descriptor-plus-fingerprint baseline was small but statistically significant in the paired test (ΔAUROC = 0.0096; p = 0.0425). Scaffold-grouped cross-validation did not show a consistent ranking advantage from adding the Stage 1 scores. On strict B3DB, MechBBB achieved an AUROC of 0.894 and an AUPRC of 0.893, with no significant ranking advantage over the baseline. On the matched B3DB subset (n = 4043), MechBBB achieved higher accuracy (0.814 vs. 0.682), balanced accuracy (0.806 vs. 0.693), sensitivity (0.910 vs. 0.541), and MCC (0.631 vs. 0.401) than SwissADME BOILED-Egg, whereas BOILED-Egg had higher specificity (0.845 vs. 0.703). SHAP analysis identified TPSA, NumHDonors, and p_pampa among the leading contributors. Conclusions: MechBBB provides a leakage controlled and externally tested framework that combines competitive BBB permeability prediction with per-compound transport-related information. The Stage 1 scores produced a modest internal improvement but did not provide a consistent external ranking advantage, indicating that their main added value is transport-related context rather than a universal increase in predictive accuracy. MechBBB can be used to prioritize compounds for CNS drug discovery before experimental BBB testing and identify compounds that may warrant follow-up studies of passive permeability or active transport. The MechBBB web tool is freely available.

PharmaceuticalsVol. 19(10)
Georgetown University (US), University of Michigan (US), Georgetown University Medical Center (US), University of Maryland, College Park (US)
Openalex Percentile: Top 9%
Computational Drug Discovery Methods
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.