Robust annotation and discovery of novel cell types in single-cell ATAC-seq data through cross-modal reference alignment

Accurate cell type annotation is essential for revealing the dynamic, cell type-specific accessibility of regulatory elements from single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq) data. However, unlike the more mature single-cell RNA-sequencing (scRNA-seq) cell type annotation workflows, scATAC-seq cell type annotation remains challenging due to extreme sparsity, high dimensionality, the scarcity of labelled scATAC references, and pronounced batch effects across datasets. To enable annotation without relying on extensive scATAC labels, we introduce CARA, a cross-omics Bayesian framework that transfers cell type knowledge from scRNA-seq to scATAC-seq. CARA simultaneously enables cell type annotation, batch correction, and detection of cell types absent from the reference by combining pretraining on scRNA-seq data and semisupervised learning on target scATAC-seq data, along with distribution alignment, dynamic class weighting, and novel cell type detection. Across diverse benchmark datasets, CARA consistently outperforms baseline methods, remaining robust to batch effects. CARA's cross-modal alignment and robust annotation strategy preserve fine-grained lineage structure, enabling reconstruction of the hematopoietic differentiation trajectory. Through multidimensional uncertainty and latent-space clustering, CARA identifies novel, rare, or aberrant populations outside the reference cell type space, providing candidates for further biological validation and perturbation. Using an omics-specific generative framework and distribution alignment, CARA delivers accurate knowledge transfer and detects novel cell types in single-cell DNA methylation data, demonstrating seamless extensibility to new modalities. Ultimately, CARA offers a powerful and flexible solution for cross-modal cell type annotation in complex single-cell settings, facilitating the discovery of novel cell types and mechanistic insight into cell type-specific regulation across diverse analyses.

Authors

Institutions

Publication Details

Journal
Genome Research
Published
2026-08-28
DOI
https://doi.org/10.1101/gr.281981.126
Primary Topic
Single-cell and spatial transcriptomics
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Robust annotation and discovery of novel cell types in single-cell ATAC-seq data through cross-modal reference alignment

Wenhao Zhang, Shengquan Chen, Ying Wang, Lan Cao et al.
Genome Research
Single-cell and spatial transcriptomics
preprint

Robust annotation and discovery of novel cell types in single-cell ATAC-seq data through cross-modal reference alignment

Wenhao Zhang, Shengquan Chen, Ying Wang, Lan Cao, Yushuang He, Yongyu Long, Feng Zhou
preprint en

Abstract

Accurate cell type annotation is essential for revealing the dynamic, cell type-specific accessibility of regulatory elements from single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq) data. However, unlike the more mature single-cell RNA-sequencing (scRNA-seq) cell type annotation workflows, scATAC-seq cell type annotation remains challenging due to extreme sparsity, high dimensionality, the scarcity of labelled scATAC references, and pronounced batch effects across datasets. To enable annotation without relying on extensive scATAC labels, we introduce CARA, a cross-omics Bayesian framework that transfers cell type knowledge from scRNA-seq to scATAC-seq. CARA simultaneously enables cell type annotation, batch correction, and detection of cell types absent from the reference by combining pretraining on scRNA-seq data and semisupervised learning on target scATAC-seq data, along with distribution alignment, dynamic class weighting, and novel cell type detection. Across diverse benchmark datasets, CARA consistently outperforms baseline methods, remaining robust to batch effects. CARA's cross-modal alignment and robust annotation strategy preserve fine-grained lineage structure, enabling reconstruction of the hematopoietic differentiation trajectory. Through multidimensional uncertainty and latent-space clustering, CARA identifies novel, rare, or aberrant populations outside the reference cell type space, providing candidates for further biological validation and perturbation. Using an omics-specific generative framework and distribution alignment, CARA delivers accurate knowledge transfer and detects novel cell types in single-cell DNA methylation data, demonstrating seamless extensibility to new modalities. Ultimately, CARA offers a powerful and flexible solution for cross-modal cell type annotation in complex single-cell settings, facilitating the discovery of novel cell types and mechanistic insight into cell type-specific regulation across diverse analyses.

Genome Research
Xiamen University (CN), Nankai University (CN)
Single-cell and spatial transcriptomics
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.