Domain classification of archaeal proteomes reveals conserved fold repertoire

Archaea represent one of the three domains of cellular life and yet account for fewer than 1% of experimentally determined protein structures, leaving the extent of their structural novelty unknown. Here we present a systematic domain-level classification of 124,075 proteins from 65 archaeal classes spanning 21 phyla and all major lineages, using both AFDB and newly predicted AlphaFold3 structures classified against the Evolutionary Classification of protein Domains (ECOD). Archaeal proteins span 987 of the 2,457 ECOD X-groups defined across all cellular life, roughly 40% of known fold diversity captured within a single domain of life. Clustering by Foldseek recovered structural relationships for 63% of domains that are singletons by sequence comparison. To characterize the 21% of proteins lacking high-confidence classification, we applied successive filters for structure prediction confidence, protein length, and structural cluster context, reducing 8,452 domain-free proteins to a small number of well-folded structural orphans (less than 0.1% of the dataset). The unclassified fraction is dominated by sub-threshold matches (matches below the 0.85 DPAM confidence cutoff for high-confidence T-group assignment) to known folds (14% of all proteins) and low-confidence structure predictions (5%), not by novel structures. These results demonstrate that the protein fold repertoire at the single-domain level is broadly conserved across the deepest phylogenetic distances in cellular life, and that the gap between archaeal and well-characterized proteomes reflects classification sensitivity for divergent sequences rather than unexplored structural diversity.

Authors

Institutions

Publication Details

Journal
PLoS Computational Biology
Published
2026-09-18
DOI
https://doi.org/10.1371/journal.pcbi.1014188
Primary Topic
Machine Learning in Bioinformatics
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Domain classification of archaeal proteomes reveals conserved fold repertoire

R. Dustin Schaeffer, Jimin Pei, Kirill E. Medvedev, Rui Guo et al.
PLoS Computational Biology
Machine Learning in Bioinformatics
article

Domain classification of archaeal proteomes reveals conserved fold repertoire

R. Dustin Schaeffer, Jimin Pei, Kirill E. Medvedev, Rui Guo, Qian Cong, Nick Grishin, Jing Zhang
article en

Abstract

Archaea represent one of the three domains of cellular life and yet account for fewer than 1% of experimentally determined protein structures, leaving the extent of their structural novelty unknown. Here we present a systematic domain-level classification of 124,075 proteins from 65 archaeal classes spanning 21 phyla and all major lineages, using both AFDB and newly predicted AlphaFold3 structures classified against the Evolutionary Classification of protein Domains (ECOD). Archaeal proteins span 987 of the 2,457 ECOD X-groups defined across all cellular life, roughly 40% of known fold diversity captured within a single domain of life. Clustering by Foldseek recovered structural relationships for 63% of domains that are singletons by sequence comparison. To characterize the 21% of proteins lacking high-confidence classification, we applied successive filters for structure prediction confidence, protein length, and structural cluster context, reducing 8,452 domain-free proteins to a small number of well-folded structural orphans (less than 0.1% of the dataset). The unclassified fraction is dominated by sub-threshold matches (matches below the 0.85 DPAM confidence cutoff for high-confidence T-group assignment) to known folds (14% of all proteins) and low-confidence structure predictions (5%), not by novel structures. These results demonstrate that the protein fold repertoire at the single-domain level is broadly conserved across the deepest phylogenetic distances in cellular life, and that the gap between archaeal and well-characterized proteomes reflects classification sensitivity for divergent sequences rather than unexplored structural diversity.

PLoS Computational BiologyVol. 22(9)
University of Central Florida (US), Southwestern Medical Center (US), Institut thématique Génétique, génomique et bioinformatique (FR), Southwestern Medical Center (US), The University of Texas Southwestern Medical Center (US)
Welch Foundation, National Institute of General Medical Sciences, National Institute of Allergy and Infectious Diseases
Openalex Percentile: Top 18%
Machine Learning in Bioinformatics
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.