Predicting Protein Aggregation Using Artificial Intelligence
Artificial intelligence is reshaping how researchers predict protein aggregation, a process central to biopharmaceutical stability and neurodegenerative disease. This review synthesizes recent machine learning and deep learning approaches to aggregation prediction, tracing the field from classical sequence based descriptors to protein language models and structure aware, multimodal architectures. It examines how AI integrates sequence composition, hydrophobicity, secondary structure, and AlphaFold derived structural information to capture nonlinear determinants of misfolding, nucleation, and amyloid formation (Wang, Dai, & Zhang, 2025; Abramson et al., 2024). Deep learning models and protein language model embeddings consistently outperform handcrafted feature baselines, particularly for aggregation prone region identification (Cima et al., 2025; Eschbach, Deibler, Korani, & Swanson, 2026). However, a persistent gap separates strong benchmark performance from genuine biological and clinical utility, driven by limited high quality datasets, class imbalance, weak interpretability, and insufficient experimental validation (Bolognesi et al., 2025). The review concludes that closing this gap requires standardized benchmarks, explainable AI methods, and stronger integration between computational prediction and experimental biology, positioning AI as a complement to, rather than a replacement for, laboratory validation.
Authors
- Mariam Fatima (ORCID: https://orcid.org/0009-0008-6355-4269)
Institutions
- University of Agriculture Faisalabad (PK)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-08-26
- DOI
- https://doi.org/10.5281/zenodo.22106464
- Primary Topic
- Machine Learning in Bioinformatics
- Type
- article
- Field-Weighted Citation Impact
- 0.00