Evaluation of the Complementarity of Three Major Databases of Protein–Ligand Binding Affinity Data

Abstract Protein–ligand binding affinity data underpin a variety of tasks at the early stage of drug discovery. Public databases including ChEMBL, BindingDB, and PDBbind curate such data with distinct focuses, yet their data overlap and complementary attributes remain insufficiently quantified. In this work, we systematically cross-mapped protein–ligand binding data in PDBbind (version 2024) against data records sourced from ChEMBL (version 35) and BindingDB (version 2025.08). Our analysis revealed that merely 22.1% and 23.9% of PDBbind’s annotated protein–ligand complexes possess matching binding data records in ChEMBL and BindingDB, respectively, confirming that the PDBbind dataset is not a subset of the other two larger repositories. This limited overlap arises from fundamental differences in scope: ChEMBL and BindingDB center on drug discovery datasets, whereas the data entries in PDBbind encompass a broader chemical and biological landscape. Our analysis further uncovered substantial data redundancy in ChEMBL and BindingDB, where 40%–50% of protein–ligand interaction pairs are associated with multiple, discrepant binding affinity values. Currently, no single resource achieves a comprehensive coverage of available protein–ligand binding data, mandating researchers to adopt resource-selection strategies customized to their goals. While ChEMBL and BindingDB host extensive bioactivity archives, PDBbind directly links three-dimensional complex structures with carefully curated binding affinity data. This unique value makes it best suited for structure-based studies.

Authors

Institutions

Publication Details

Journal
Journal of Chemical Information and Modeling
Published
2026-09-29
DOI
https://doi.org/10.1021/acs.jcim.6c02172
Primary Topic
Computational Drug Discovery Methods
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Evaluation of the Complementarity of Three Major Databases of Protein–Ligand Binding Affinity Data

Renxiao Wang, Jingyuan Li, Haotian Gao, Xinchong Chen et al.
Journal of Chemical Information and Modeling
Computational Drug Discovery Methods
article

Evaluation of the Complementarity of Three Major Databases of Protein–Ligand Binding Affinity Data

Renxiao Wang, Jingyuan Li, Haotian Gao, Xinchong Chen, Yanbei Li, Yan Li, Yifei Qi, Zhen Duan, Kexin Wu, Fengfei Miao
article en

Abstract

Abstract Protein–ligand binding affinity data underpin a variety of tasks at the early stage of drug discovery. Public databases including ChEMBL, BindingDB, and PDBbind curate such data with distinct focuses, yet their data overlap and complementary attributes remain insufficiently quantified. In this work, we systematically cross-mapped protein–ligand binding data in PDBbind (version 2024) against data records sourced from ChEMBL (version 35) and BindingDB (version 2025.08). Our analysis revealed that merely 22.1% and 23.9% of PDBbind’s annotated protein–ligand complexes possess matching binding data records in ChEMBL and BindingDB, respectively, confirming that the PDBbind dataset is not a subset of the other two larger repositories. This limited overlap arises from fundamental differences in scope: ChEMBL and BindingDB center on drug discovery datasets, whereas the data entries in PDBbind encompass a broader chemical and biological landscape. Our analysis further uncovered substantial data redundancy in ChEMBL and BindingDB, where 40%–50% of protein–ligand interaction pairs are associated with multiple, discrepant binding affinity values. Currently, no single resource achieves a comprehensive coverage of available protein–ligand binding data, mandating researchers to adopt resource-selection strategies customized to their goals. While ChEMBL and BindingDB host extensive bioactivity archives, PDBbind directly links three-dimensional complex structures with carefully curated binding affinity data. This unique value makes it best suited for structure-based studies.

Journal of Chemical Information and Modeling
Fudan University (CN), Shanghai CASB Biotechnology (China) (CN)
Openalex Percentile: Top 9%
Computational Drug Discovery Methods
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.