UIDDA: a unified-input model-classifier combination framework for drug-disease association prediction

Drug-disease association (DDA) prediction methods differ substantially in data preprocessing, feature construction, model architecture, negative-sample treatment, and evaluation protocols, which limits fair comparison across studies. We developed UIDDA, a unified-input model-classifier benchmark that standardizes data processing and separates representation learning from downstream classification. UIDDA evaluates nine feature extraction models combined with four classification heads on four public DDA datasets. To prevent information leakage, association matrices, Gaussian interaction profile similarities, heterogeneous biomedical networks, and association-dependent representations were reconstructed independently within each training fold. Under the pair-level Random-U setting, FD-MSGL and AMDGT achieved the strongest overall performance, whereas DeepDR, HNetDNN, and LDSCNN showed lower average discrimination. Model selection produced substantially larger performance variation than classifier selection; the mean AUC difference between the best and worst models was 0.266, compared with 0.032 among classification heads. Dataset difficulty also varied, with the B-dataset achieving the highest overall mean AUC and the sparse T-dataset the lowest. The benchmark remained informative under alternative treatments of unobserved pairs: overall AUC/AUPR values were 0.724/0.732 for Random-U, 0.692/0.709 for Hard-U, and 0.675/0.694 for nnPU. To assess entity-level generalization, we considered three entity-held-out settings: drug-held-out, disease-held-out, and both-held-out, in which the test drugs, test diseases, or both entity types were excluded from training, respectively. The corresponding mean AUC values were 0.692, 0.614, and 0.552, showing that generalization became more difficult when unseen entities were introduced. Adding MolVis-inspired three-dimensional molecular features to FD-MSGL produced only small and dataset-dependent changes. A frozen top-10 retrospective literature assessment provided a limited qualitative plausibility check for ranked model outputs. UIDDA provides a standardized, leakage-controlled, and reproducible framework for separating the impact of feature representation and classification decisions and for comparing and evaluating DDA prediction methods. The results indicate that model selection is the main source of performance variation, and the conclusions depend on the underlying processing and evaluation settings.

Authors

Institutions

Publication Details

Journal
BMC Bioinformatics
Published
2026-08-26
DOI
https://doi.org/10.1186/s12859-026-06626-6
Primary Topic
Machine Learning in Healthcare
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

UIDDA: a unified-input model-classifier combination framework for drug-disease association prediction

Tingting Guo, Dayu Zou, Huirui Han, Jian Chen et al.
BMC Bioinformatics
Machine Learning in Healthcare
article

UIDDA: a unified-input model-classifier combination framework for drug-disease association prediction

Tingting Guo, Dayu Zou, Huirui Han, Jian Chen, Xingjun Cai, Zhengxin Chen, Limei Wang, Xinying Liu, Jin Li
article en

Abstract

Drug-disease association (DDA) prediction methods differ substantially in data preprocessing, feature construction, model architecture, negative-sample treatment, and evaluation protocols, which limits fair comparison across studies. We developed UIDDA, a unified-input model-classifier benchmark that standardizes data processing and separates representation learning from downstream classification. UIDDA evaluates nine feature extraction models combined with four classification heads on four public DDA datasets. To prevent information leakage, association matrices, Gaussian interaction profile similarities, heterogeneous biomedical networks, and association-dependent representations were reconstructed independently within each training fold. Under the pair-level Random-U setting, FD-MSGL and AMDGT achieved the strongest overall performance, whereas DeepDR, HNetDNN, and LDSCNN showed lower average discrimination. Model selection produced substantially larger performance variation than classifier selection; the mean AUC difference between the best and worst models was 0.266, compared with 0.032 among classification heads. Dataset difficulty also varied, with the B-dataset achieving the highest overall mean AUC and the sparse T-dataset the lowest. The benchmark remained informative under alternative treatments of unobserved pairs: overall AUC/AUPR values were 0.724/0.732 for Random-U, 0.692/0.709 for Hard-U, and 0.675/0.694 for nnPU. To assess entity-level generalization, we considered three entity-held-out settings: drug-held-out, disease-held-out, and both-held-out, in which the test drugs, test diseases, or both entity types were excluded from training, respectively. The corresponding mean AUC values were 0.692, 0.614, and 0.552, showing that generalization became more difficult when unseen entities were introduced. Adding MolVis-inspired three-dimensional molecular features to FD-MSGL produced only small and dataset-dependent changes. A frozen top-10 retrospective literature assessment provided a limited qualitative plausibility check for ranked model outputs. UIDDA provides a standardized, leakage-controlled, and reproducible framework for separating the impact of feature representation and classification decisions and for comparing and evaluating DDA prediction methods. The results indicate that model selection is the main source of performance variation, and the conclusions depend on the underlying processing and evaluation settings.

BMC Bioinformatics
Hainan Medical University (CN)
National Natural Science Foundation of China, Natural Science Foundation of Hainan Province
Peace, Justice and strong institutions
Openalex Percentile: Top 8%
Machine Learning in Healthcare
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.