scTrans-SVD: dual-view learning with a low-rank representation for label-limited cell type annotation in single-cell transcriptomics

Single-cell transcriptomic profiling reveals cellular heterogeneity, but reliable cell type labels are often available for only a subset of profiles. This setting motivates approaches that can use label-withheld training expression while retaining measured gene expression as a direct predictive input. We developed scTrans-SVD, a split-specific dual-view framework for label-limited cell type annotation. One view retains measured gene expression, while the other provides a low-rank representation learned only from model-training profiles by SVD-initialized pseudo-Huber matrix factorization. Aligned gene-subvector tokens from the two views are encoded with shared weights and combined by fixed equal-weight logit fusion. We evaluated scTrans-SVD on held-out profiles from 13 public single-cell transcriptomic datasets using five fixed seeds. Under a class-aware 10% total-label-budget protocol (realized dataset means, 10.01–11.39%), scTrans-SVD achieved mean Macro-F1 0.8820. Aggregate values for the seven comparator methods ranged from 0.2785 to 0.8442, and scTrans-SVD had a higher dataset mean than all seven comparator methods on 9 of 13 datasets. Independently trained measured-input-only and MF-reconstructed-input-only models achieved 0.8637 and 0.8665, respectively. The complete dual-view model was higher than each single-view model on 12 of 13 datasets and was higher than both on 11 datasets. Group-held-out, pancreas reference-query and paired protein-modality analyses further characterized performance and scope. scTrans-SVD combines measured expression with a split-specific low-rank view to use structure in label-withheld training profiles without allowing held-out profiles to influence factorization. The benchmark and independent single-view patterns were consistent with a potential benefit from combining the two views under limited label access. These findings suggest that the framework may be considered for closed-set transcriptomic annotation when profiles are abundant but trusted labels are limited, while broader cross-study and unseen-cell-type evaluation remain important directions.

Authors

Institutions

Publication Details

Journal
BioData Mining
Published
2026-09-18
DOI
https://doi.org/10.1186/s13040-026-00602-9
Primary Topic
Single-cell and spatial transcriptomics
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

scTrans-SVD: dual-view learning with a low-rank representation for label-limited cell type annotation in single-cell transcriptomics

Juseong Kim, Giltae Song, Sanghun Sel, Jason Hilton
BioData Mining
Single-cell and spatial transcriptomics
article

scTrans-SVD: dual-view learning with a low-rank representation for label-limited cell type annotation in single-cell transcriptomics

Juseong Kim, Giltae Song, Sanghun Sel, Jason Hilton
article en

Abstract

Single-cell transcriptomic profiling reveals cellular heterogeneity, but reliable cell type labels are often available for only a subset of profiles. This setting motivates approaches that can use label-withheld training expression while retaining measured gene expression as a direct predictive input. We developed scTrans-SVD, a split-specific dual-view framework for label-limited cell type annotation. One view retains measured gene expression, while the other provides a low-rank representation learned only from model-training profiles by SVD-initialized pseudo-Huber matrix factorization. Aligned gene-subvector tokens from the two views are encoded with shared weights and combined by fixed equal-weight logit fusion. We evaluated scTrans-SVD on held-out profiles from 13 public single-cell transcriptomic datasets using five fixed seeds. Under a class-aware 10% total-label-budget protocol (realized dataset means, 10.01–11.39%), scTrans-SVD achieved mean Macro-F1 0.8820. Aggregate values for the seven comparator methods ranged from 0.2785 to 0.8442, and scTrans-SVD had a higher dataset mean than all seven comparator methods on 9 of 13 datasets. Independently trained measured-input-only and MF-reconstructed-input-only models achieved 0.8637 and 0.8665, respectively. The complete dual-view model was higher than each single-view model on 12 of 13 datasets and was higher than both on 11 datasets. Group-held-out, pancreas reference-query and paired protein-modality analyses further characterized performance and scope. scTrans-SVD combines measured expression with a split-specific low-rank view to use structure in label-withheld training profiles without allowing held-out profiles to influence factorization. The benchmark and independent single-view patterns were consistent with a potential benefit from combining the two views under limited label access. These findings suggest that the framework may be considered for closed-set transcriptomic annotation when profiles are abundant but trusted labels are limited, while broader cross-study and unseen-cell-type evaluation remain important directions.

BioData Mining
Pusan National University (KR), Stanford University (US)
Openalex Percentile: Top 18%
Single-cell and spatial transcriptomics
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.