TCM-Complexity: a complexity-aware hybrid active learning framework for drug–target interaction prediction
Drug–target interaction (DTI) prediction is an essential task in computer-aided drug discovery and pharmacological mechanism elucidation. High-quality DTI annotations typically rely on costly wet-lab experiments, and limited experimental resources make it difficult to obtain large-scale, reliable labeled datasets. Therefore, effectively assessing the utility of candidate samples and improving data utilization efficiency under limited annotation budgets remain critical challenges. Active learning offers a promising solution to reduce annotation costs; however, existing methods mainly focus on model prediction uncertainty or sample distribution characteristics and do not fully exploit the latent structural information embedded in complex biological representation spaces. To address these limitations, we propose TCM-Complexity, a complexity-aware and stage-adaptive active learning framework for DTI prediction. The framework constructs a joint drug–target embedding space to characterize the local structural complexity of candidate samples and incorporates this information as a novel sample selection criterion. By integrating model prediction information with data structural characteristics, a two-stage exploration–mining strategy is developed to dynamically optimize sample selection across different learning stages. Experimental evaluations on three public DTI datasets, namely DAVIS, BioSNAP, and TCMSP, showed that TCM-Complexity achieved the best overall performance across most evaluated metrics in the primary random-split experiments under a labeling budget corresponding to 20% of the training pool. Compared with the strongest baseline, TCM, the maximum observed improvements were 1.87% in ROC_AUC and 1.79% in PR_AUC. The proposed method also achieved more than 90% of the predictive performance of the fully supervised models across the evaluated settings, with the maximum relative performance approaching 98%. Additional cold-drug, cold-target, and multi-seed analyses indicated that the magnitude of the performance advantage was dataset- and metric-dependent. TCM-Complexity can improve DTI prediction performance and sample utilization efficiency under the evaluated limited-annotation settings. By jointly considering model prediction information and the local structural complexity of candidate samples, the proposed framework provides a complementary structure-aware active-learning strategy for computational drug discovery under constrained annotation resources.
Authors
- Zhaoxing Xu (ORCID: https://orcid.org/0009-0007-0923-0063)
- Wangping Xiong (ORCID: https://orcid.org/0000-0003-4992-8558)
- Yi Zhang (ORCID: https://orcid.org/0000-0002-7453-6188)
- Xin Cheng
- Xin Cheng
- Delong Yuan
- Pinzheng Liu
Institutions
- Beijing Institute of Fashion Technology (CN)
- Jiangxi University of Traditional Chinese Medicine (CN)
- Jiangxi University of Water Resources and Electric Power (CN)
Publication Details
- Journal
- BMC Bioinformatics
- Published
- 2026-10-03
- DOI
- https://doi.org/10.1186/s12859-026-06684-w
- Primary Topic
- Computational Drug Discovery Methods
- Type
- article
- Field-Weighted Citation Impact
- 0.00