Benchmarking descriptor-based AI approaches for predicting the BBB permeability of drug-like xenobiotics
Accurate prediction of blood-brain barrier (BBB) permeability is widely recognized as important for drug discovery, neurotoxicity assessment, and evaluating xenobiotic exposure linked to neurodegenerative diseases. In silico models, particularly those based on artificial intelligence (AI), are increasingly used to predict small-molecule BBB permeability. However, challenges remain due to heterogeneous datasets, limited sample sizes, and the incomplete capture of relevant molecular features, which affect model reliability and interpretability. Feature selection plays a key role in enhancing the robustness and reproducibility of model predictions. This study tests a workflow combining ensemble feature selection, which integrates multiple feature selection methods, with different machine-learning based algorithms to identify a compact and interpretable model for BBB permeability prediction. Physicochemical and structural descriptors computed with RDKit, AlvaDesc, Mordred, and MACCS fingerprints were benchmarked in terms of model performance. A dataset of 8,159 small molecules from four literature and commercial sources was curated by harmonizing labels and applying molecular standardization. By combining adaptive correlation filtering with ensemble feature selection across diverse molecular descriptors, we derived a compact ten-feature Random Forest model achieving 0.83 balanced accuracy in nested cross-validation. On an independent FDA-approved drug set, the model achieved a balanced accuracy of 0.77, indicating competitive performance within the specific external validation setting considered here. This study provides a systematic benchmark integrating molecular geometry and protonation-state optimization, expert-guided dataset curation, molecular representations derived from RDKit, AlvaDesc, Mordred, and MACCS fingerprints, ensemble feature selection, and six machine-learning algorithms for BBB permeability prediction. Unlike approaches based on individual descriptor sets or high-dimensional representations, the proposed workflow explicitly considers the trade-off between predictive performance and model complexity, identifying a chemically interpretable 10-feature Random Forest model that retains competitive discriminative performance and was evaluated on an independent set of 27 FDA-approved drugs.
Authors
- Francesco Lopresti (ORCID: https://orcid.org/0000-0003-3893-0820)
- Claudia Coronnello (ORCID: https://orcid.org/0000-0001-5962-8642)
- Nicolina Sciaraffa (ORCID: https://orcid.org/0000-0002-3380-0181)
- Maria Rita Gulotta (ORCID: https://orcid.org/0000-0001-5032-4888)
- Ugo Perricone (ORCID: https://orcid.org/0000-0002-2181-2468)
- Martina Valentino
Institutions
- Ri.MED (IT)
- University of Palermo (IT)
Publication Details
- Journal
- Journal of Cheminformatics
- Published
- 2026-09-25
- DOI
- https://doi.org/10.1186/s13321-026-01309-z
- Primary Topic
- Computational Drug Discovery Methods
- Type
- article
- Field-Weighted Citation Impact
- 0.00