Temporally Validated Models for Target “Tclin-likeness”

Abstract Identifying which human proteins are likely to become validated drug targets remains a challenge in drug discovery. We build on the “Tclin” concept, which encompasses proteins for which at least one approved drug exerts its mode of action (MoA); these comprise only ∼3% of the human proteome. We developed disease-agnostic, knowledge-graph machine learning (KGML) models that estimate “Tclin-likeness”: how closely a non-Tclin protein resembles established MoA drug targets in knowledge-graph feature space. Single-class support vector machines (1SVM) and XGBoost classifiers were trained on Pharos data current as of December 2017, using static (Gene Ontology, Reactome, the Kyoto encyclopedia for genes and genomes, STRING, HPA) and dynamic (genotype-tissue expression, LINCS, DISEASES, cancer cell line encyclopedia) features. To prevent data leakage, 115 proteins that acquired Tclin status after January 2018 were withheld entirely from training. Sensitivity on this external set was 88%–95% across individual models. Because non-Tclin proteins are unlabeled rather than confirmed negatives, specificity and precision are reported as bounds: balanced accuracy is 54%–75% for individual models and 73.7% for the eight-model consensus, which flags 5019 non-Tclin proteins (27.4%). The 1SVM models also identified a set of 1290 consensus negatives for XGBoost training. Benchmarked against two KGML druggability models, DrugnomeAI and PINNED, a consensus of 1133 proteins emerged. The Tclin-like consensus recovered 145 of 181 non-Tclin proteins in Phase 1–3 clinical trials (80.1%) and 23 of 28 proteins that reached Tclin status in 2023–2025 (82.1%), each approximately 3-fold enriched over background; it is statistically indistinguishable from PINNED (McNemar p = 0.42) while flagging 40% fewer proteins and outperforms DrugnomeAI (p = 0.0015). The three-model consensus places the druggable genome between 1860 and 5745 genes. Tclin-likeness is a prior, not a proof of tractability: these features encode nothing about binding pockets, ligandability or therapeutic modality. The “Tclin-like” consensus score is best used as an independent filter alongside disease-specific evidence and modality-specific tractability assessment.

Authors

Institutions

Publication Details

Journal
Journal of Chemical Information and Modeling
Published
2026-10-07
DOI
https://doi.org/10.1021/acs.jcim.6c03088
Primary Topic
Computational Drug Discovery Methods
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Temporally Validated Models for Target “Tclin-likeness”

Cristian Bologa, Tudor Ionel Oprea, Suman Sirimulla, Mohammed A. Quazi et al.
Journal of Chemical Information and Modeling
Computational Drug Discovery Methods
article

Temporally Validated Models for Target “Tclin-likeness”

Cristian Bologa, Tudor Ionel Oprea, Suman Sirimulla, Mohammed A. Quazi, Alexei Pushechnikov, Nikolay Savchuk, Bill Farley
article en

Abstract

Abstract Identifying which human proteins are likely to become validated drug targets remains a challenge in drug discovery. We build on the “Tclin” concept, which encompasses proteins for which at least one approved drug exerts its mode of action (MoA); these comprise only ∼3% of the human proteome. We developed disease-agnostic, knowledge-graph machine learning (KGML) models that estimate “Tclin-likeness”: how closely a non-Tclin protein resembles established MoA drug targets in knowledge-graph feature space. Single-class support vector machines (1SVM) and XGBoost classifiers were trained on Pharos data current as of December 2017, using static (Gene Ontology, Reactome, the Kyoto encyclopedia for genes and genomes, STRING, HPA) and dynamic (genotype-tissue expression, LINCS, DISEASES, cancer cell line encyclopedia) features. To prevent data leakage, 115 proteins that acquired Tclin status after January 2018 were withheld entirely from training. Sensitivity on this external set was 88%–95% across individual models. Because non-Tclin proteins are unlabeled rather than confirmed negatives, specificity and precision are reported as bounds: balanced accuracy is 54%–75% for individual models and 73.7% for the eight-model consensus, which flags 5019 non-Tclin proteins (27.4%). The 1SVM models also identified a set of 1290 consensus negatives for XGBoost training. Benchmarked against two KGML druggability models, DrugnomeAI and PINNED, a consensus of 1133 proteins emerged. The Tclin-like consensus recovered 145 of 181 non-Tclin proteins in Phase 1–3 clinical trials (80.1%) and 23 of 28 proteins that reached Tclin status in 2023–2025 (82.1%), each approximately 3-fold enriched over background; it is statistically indistinguishable from PINNED (McNemar p = 0.42) while flagging 40% fewer proteins and outperforms DrugnomeAI (p = 0.0015). The three-model consensus places the druggable genome between 1860 and 5745 genes. Tclin-likeness is a prior, not a proof of tractability: these features encode nothing about binding pockets, ligandability or therapeutic modality. The “Tclin-like” consensus score is best used as an independent filter alongside disease-specific evidence and modality-specific tractability assessment.

Journal of Chemical Information and Modeling
West Virginia University (US), University of New Mexico (US)
Openalex Percentile: Top 13%
Computational Drug Discovery Methods
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.