TCM-Complexity: a complexity-aware hybrid active learning framework for drug–target interaction prediction

Drug–target interaction (DTI) prediction is an essential task in computer-aided drug discovery and pharmacological mechanism elucidation. High-quality DTI annotations typically rely on costly wet-lab experiments, and limited experimental resources make it difficult to obtain large-scale, reliable labeled datasets. Therefore, effectively assessing the utility of candidate samples and improving data utilization efficiency under limited annotation budgets remain critical challenges. Active learning offers a promising solution to reduce annotation costs; however, existing methods mainly focus on model prediction uncertainty or sample distribution characteristics and do not fully exploit the latent structural information embedded in complex biological representation spaces. To address these limitations, we propose TCM-Complexity, a complexity-aware and stage-adaptive active learning framework for DTI prediction. The framework constructs a joint drug–target embedding space to characterize the local structural complexity of candidate samples and incorporates this information as a novel sample selection criterion. By integrating model prediction information with data structural characteristics, a two-stage exploration–mining strategy is developed to dynamically optimize sample selection across different learning stages. Experimental evaluations on three public DTI datasets, namely DAVIS, BioSNAP, and TCMSP, showed that TCM-Complexity achieved the best overall performance across most evaluated metrics in the primary random-split experiments under a labeling budget corresponding to 20% of the training pool. Compared with the strongest baseline, TCM, the maximum observed improvements were 1.87% in ROC_AUC and 1.79% in PR_AUC. The proposed method also achieved more than 90% of the predictive performance of the fully supervised models across the evaluated settings, with the maximum relative performance approaching 98%. Additional cold-drug, cold-target, and multi-seed analyses indicated that the magnitude of the performance advantage was dataset- and metric-dependent. TCM-Complexity can improve DTI prediction performance and sample utilization efficiency under the evaluated limited-annotation settings. By jointly considering model prediction information and the local structural complexity of candidate samples, the proposed framework provides a complementary structure-aware active-learning strategy for computational drug discovery under constrained annotation resources.

Authors

Institutions

Publication Details

Journal
BMC Bioinformatics
Published
2026-10-03
DOI
https://doi.org/10.1186/s12859-026-06684-w
Primary Topic
Computational Drug Discovery Methods
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

TCM-Complexity: a complexity-aware hybrid active learning framework for drug–target interaction prediction

Zhaoxing Xu, Wangping Xiong, Yi Zhang, Xin Cheng et al.
BMC Bioinformatics
Computational Drug Discovery Methods
article

TCM-Complexity: a complexity-aware hybrid active learning framework for drug–target interaction prediction

Zhaoxing Xu, Wangping Xiong, Yi Zhang, Xin Cheng, Xin Cheng, Delong Yuan, Pinzheng Liu
article en

Abstract

Drug–target interaction (DTI) prediction is an essential task in computer-aided drug discovery and pharmacological mechanism elucidation. High-quality DTI annotations typically rely on costly wet-lab experiments, and limited experimental resources make it difficult to obtain large-scale, reliable labeled datasets. Therefore, effectively assessing the utility of candidate samples and improving data utilization efficiency under limited annotation budgets remain critical challenges. Active learning offers a promising solution to reduce annotation costs; however, existing methods mainly focus on model prediction uncertainty or sample distribution characteristics and do not fully exploit the latent structural information embedded in complex biological representation spaces. To address these limitations, we propose TCM-Complexity, a complexity-aware and stage-adaptive active learning framework for DTI prediction. The framework constructs a joint drug–target embedding space to characterize the local structural complexity of candidate samples and incorporates this information as a novel sample selection criterion. By integrating model prediction information with data structural characteristics, a two-stage exploration–mining strategy is developed to dynamically optimize sample selection across different learning stages. Experimental evaluations on three public DTI datasets, namely DAVIS, BioSNAP, and TCMSP, showed that TCM-Complexity achieved the best overall performance across most evaluated metrics in the primary random-split experiments under a labeling budget corresponding to 20% of the training pool. Compared with the strongest baseline, TCM, the maximum observed improvements were 1.87% in ROC_AUC and 1.79% in PR_AUC. The proposed method also achieved more than 90% of the predictive performance of the fully supervised models across the evaluated settings, with the maximum relative performance approaching 98%. Additional cold-drug, cold-target, and multi-seed analyses indicated that the magnitude of the performance advantage was dataset- and metric-dependent. TCM-Complexity can improve DTI prediction performance and sample utilization efficiency under the evaluated limited-annotation settings. By jointly considering model prediction information and the local structural complexity of candidate samples, the proposed framework provides a complementary structure-aware active-learning strategy for computational drug discovery under constrained annotation resources.

BMC Bioinformatics
Beijing Institute of Fashion Technology (CN), Jiangxi University of Traditional Chinese Medicine (CN), Jiangxi University of Water Resources and Electric Power (CN)
Openalex Percentile: Top 9%
Computational Drug Discovery Methods
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.