Seq2DomML: Protein domain boundary prediction from bacterial protein sequences using a machine learning framework

Abstract Purpose The identification of protein domain boundaries is important for designing stable single-domain constructs for protein structural solution. For proteins without experimentally determined structures, boundary selection often relies on conserved domain annotations, predicted structures, or iterative truncation screening. Seq2DomML is developed as a sequence-based machine learning framework for predicting residue-level domain partitions in bacterial proteins without requiring structural coordinates or homologous domain assignments as input. Methods Structural domain annotations from the Evolutionary Classification of Domains (ECOD) database were converted into binary residue labels, including residues within annotated domains and domain boundaries, linkers, and terminal flanking regions. Each residue was initially encoded using 2389 amino acid identity, positional, physicochemical, and interaction features. The feature space was reduced through correlation-based pruning and multi-objective evolutionary feature selection, and the selected features were used to train a bidirectional long short-term memory (LSTM) model. Results Seq2DomML was evaluated on a held-out test set after sequences sharing at least 35% global sequence identity with the training set were removed and was compared with established tools, including InterPro, Pfam, and NCBI CDD. The classification metrics showed that Seq2DomML outperformed the comparison tools, achieving a macro F1-score of 0.734 and a within-domain F1-score of 0.924. Its domain count mean absolute error was 0.952, the second-lowest value after Pfam. Seq2DomML additionally achieved the highest mean domain Intersection over Union (IoU) of 0.846, indicating the strongest spatial agreement with the ECOD-derived domain regions among the evaluated methods. Conclusion Seq2DomML demonstrates that sequence-derived residue features can support prediction of structurally defined domain partitions without direct structural input. The model provides a complementary approach to conserved domain annotation tools and may help prioritize candidate splice regions for construct design, particularly when structural or homologous evidence is limited.

Authors

Institutions

Publication Details

Journal
BMC Bioinformatics
Published
2026-10-06
DOI
https://doi.org/10.1186/s12859-026-06666-y
Primary Topic
Machine Learning in Bioinformatics
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Seq2DomML: Protein domain boundary prediction from bacterial protein sequences using a machine learning framework

Michael D. L. Suits, Azam Asilian Bidgoli, Alana Monks
BMC Bioinformatics
Machine Learning in Bioinformatics
article

Seq2DomML: Protein domain boundary prediction from bacterial protein sequences using a machine learning framework

Michael D. L. Suits, Azam Asilian Bidgoli, Alana Monks
article en

Abstract

Abstract Purpose The identification of protein domain boundaries is important for designing stable single-domain constructs for protein structural solution. For proteins without experimentally determined structures, boundary selection often relies on conserved domain annotations, predicted structures, or iterative truncation screening. Seq2DomML is developed as a sequence-based machine learning framework for predicting residue-level domain partitions in bacterial proteins without requiring structural coordinates or homologous domain assignments as input. Methods Structural domain annotations from the Evolutionary Classification of Domains (ECOD) database were converted into binary residue labels, including residues within annotated domains and domain boundaries, linkers, and terminal flanking regions. Each residue was initially encoded using 2389 amino acid identity, positional, physicochemical, and interaction features. The feature space was reduced through correlation-based pruning and multi-objective evolutionary feature selection, and the selected features were used to train a bidirectional long short-term memory (LSTM) model. Results Seq2DomML was evaluated on a held-out test set after sequences sharing at least 35% global sequence identity with the training set were removed and was compared with established tools, including InterPro, Pfam, and NCBI CDD. The classification metrics showed that Seq2DomML outperformed the comparison tools, achieving a macro F1-score of 0.734 and a within-domain F1-score of 0.924. Its domain count mean absolute error was 0.952, the second-lowest value after Pfam. Seq2DomML additionally achieved the highest mean domain Intersection over Union (IoU) of 0.846, indicating the strongest spatial agreement with the ECOD-derived domain regions among the evaluated methods. Conclusion Seq2DomML demonstrates that sequence-derived residue features can support prediction of structurally defined domain partitions without direct structural input. The model provides a complementary approach to conserved domain annotation tools and may help prioritize candidate splice regions for construct design, particularly when structural or homologous evidence is limited.

BMC Bioinformatics
Wilfrid Laurier University (CA)
Openalex Percentile: Top 22%
Machine Learning in Bioinformatics
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.