VAEBAC: A representation-learning framework for proteome-scale prediction of amyloidogenic proteins and functional stratification of nucleic-acid-binding proteins

Abstract Motivation Nucleic-acid-binding proteins (NABPs) govern transcription, replication and genome organization, yet the distribution of aggregation susceptibility within regulatory proteomes remains unresolved. Here, we systematically examine amyloidogenic potential across the nucleic-acid-binding proteomes of four phylogenetically diverse organisms—Escherichia coli, Bacillus subtilis, Saccharomyces cerevisiae and Homo sapiens—spanning prokaryotes and eukaryotes, to determine whether aggregation-prone sequence architectures are selectively constrained within transcriptional networks. Results Using VAEBAC, a representation-learning framework that integrates large-scale amyloid-family sequences with experimentally validated annotations, we performed proteome-scale prediction of amyloidogenicity across nucleic-acid-binding proteins in each organism. Predicted amyloidogenic proteins were consistently and significantly enriched in Gene Ontology categories associated with regulation of DNA-templated transcription, RNA biosynthesis, and gene expression across all four organisms, a pattern validated using three independent annotation tools. Fisher’s exact test confirmed significant enrichment of amyloidogenic NABPs among transcription regulators in eukaryotic proteomes. Residue-level mapping demonstrated spatial separation between predicted aggregation-prone regions and DNA-binding interfaces, suggesting structural compatibility between regulatory function and conditional aggregation. Experimental validation confirmed fibrillar assembly of selected proteins. These findings support a conserved model in which aggregation susceptibility is functionally stratified across transcriptional regulatory tiers rather than uniformly suppressed within NABPs, a principle conserved across 3.5 billion years of evolution. Availability and Implementation VAEBAC is freely accessible at https://www.vaebac.com. Supplementary Information Supplementary data are available at Bioinformatics online.

Authors

Institutions

Publication Details

Journal
Bioinformatics
Published
2026-10-07
DOI
https://doi.org/10.1093/bioinformatics/btag745
Primary Topic
Protein Structure and Dynamics
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

VAEBAC: A representation-learning framework for proteome-scale prediction of amyloidogenic proteins and functional stratification of nucleic-acid-binding proteins

Owais Ahmad, Rizwan Khan, Muhammad Uzair Ashraf
Bioinformatics
Protein Structure and Dynamics
article

VAEBAC: A representation-learning framework for proteome-scale prediction of amyloidogenic proteins and functional stratification of nucleic-acid-binding proteins

Owais Ahmad, Rizwan Khan, Muhammad Uzair Ashraf
article en

Abstract

Abstract Motivation Nucleic-acid-binding proteins (NABPs) govern transcription, replication and genome organization, yet the distribution of aggregation susceptibility within regulatory proteomes remains unresolved. Here, we systematically examine amyloidogenic potential across the nucleic-acid-binding proteomes of four phylogenetically diverse organisms—Escherichia coli, Bacillus subtilis, Saccharomyces cerevisiae and Homo sapiens—spanning prokaryotes and eukaryotes, to determine whether aggregation-prone sequence architectures are selectively constrained within transcriptional networks. Results Using VAEBAC, a representation-learning framework that integrates large-scale amyloid-family sequences with experimentally validated annotations, we performed proteome-scale prediction of amyloidogenicity across nucleic-acid-binding proteins in each organism. Predicted amyloidogenic proteins were consistently and significantly enriched in Gene Ontology categories associated with regulation of DNA-templated transcription, RNA biosynthesis, and gene expression across all four organisms, a pattern validated using three independent annotation tools. Fisher’s exact test confirmed significant enrichment of amyloidogenic NABPs among transcription regulators in eukaryotic proteomes. Residue-level mapping demonstrated spatial separation between predicted aggregation-prone regions and DNA-binding interfaces, suggesting structural compatibility between regulatory function and conditional aggregation. Experimental validation confirmed fibrillar assembly of selected proteins. These findings support a conserved model in which aggregation susceptibility is functionally stratified across transcriptional regulatory tiers rather than uniformly suppressed within NABPs, a principle conserved across 3.5 billion years of evolution. Availability and Implementation VAEBAC is freely accessible at https://www.vaebac.com. Supplementary Information Supplementary data are available at Bioinformatics online.

Bioinformatics
Aligarh Muslim University (IN)
Openalex Percentile: Top 23%
Protein Structure and Dynamics
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.