A curated structural dataset of peptide–protein complexes reveals biases in existing datasets and principles of peptide binding
Peptide-protein interactions are fundamental to many biological processes, and peptide design is gaining interest due to its therapeutic potential. This has led to the emergence of various structural databases, such as PepBDB and Peptipedia, that classify peptides as polypeptides with fewer than 50 amino acids. These databases provide valuable starting points for studying peptide recognition by protein partners and are widely used for machine learning applications, docking, and scoring functions. However, under such a length definition for peptides, we have very different cases that are likely to confound the analysis of peptide binding, including miniproteins, peptide-peptide complexes, intramolecular peptide disulfide bonds, intermolecular disulfide bridges, and proteins undergoing internal cleavage, such as serpins, as well as non-natural amino acids and covalently bound cofactors. Here, we present a rigorous classification of peptide-protein complexes to generate datasets suitable for comparative energetic analysis with a focus on peptides that are unstructured in the absence of their target protein. The analysis of this dataset shows that peptide binding is typically driven by a small number of hotspot residues mainly enriched in aromatic and bulky hydrophobic side chains. Their number of hotspots and their spatial organization depend on peptide length, secondary structure, and covalent constraints. Short peptides rely on central anchor regions, whereas longer peptides distribute hotspots more broadly, with helices showing periodic spacing and β-strands relying more on backbone-mediated stabilization. Disulfide bonds further decrease the number of hotspots per peptide length by either pre-organizing the peptide or acting as covalent anchors. This work provides a curated resource and general principles for peptide recognition. It highlights the importance of structurally classifying peptide-protein complexes to avoid bias in downstream computational and machine-learning applications.
Authors
- Luís Serrano (ORCID: https://orcid.org/0000-0002-5276-1392)
- Javier Delgado (ORCID: https://orcid.org/0000-0003-1302-5445)
- Rahma Hamdani (ORCID: https://orcid.org/0009-0002-3654-0378)
Institutions
- Institució Catalana de Recerca i Estudis Avançats (ES)
- Universitat Pompeu Fabra (ES)
- Barcelona Institute of Science and Technology (ES)
- Centre for Genomic Regulation (ES)
Publication Details
- Journal
- Protein Science
- Published
- 2026-09-15
- DOI
- https://doi.org/10.1002/pro.70779
- Primary Topic
- Advanced Proteomics Techniques and Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00