SwinTransRNA: A Swin Transformer-Based Framework for Accurate Prediction of RNA Subcellular Localization
Abstract RNA subcellular localization determines the regulatory context in which microRNAs (miRNAs), circular RNAs (circRNAs), and long noncoding RNAs (lncRNAs) exert their functions, yet these RNA classes differ greatly in sequence length and circularity, making fixed-length raw-sequence representations inconvenient for cross-RNA comparison. Here, we present SwinTransRNA, a unified framework that encodes each RNA as a length-normalized 64 × 64 frequency chaos game representation (FCGR) of 6-mer composition, where every cell carries a fixed 6-mer identity and the matrix entries sum to one, enabling direct comparison across sequence lengths. A hierarchical window-attention network captures localized interactions among neighboring composition regions and enlarges the receptive field through shifted windows and patch merging. On audited benchmark data sets for nuclear/cytoplasmic lncRNA, nuclear/cytoplasmic circRNA, and intracellular/extracellular miRNA classification, SwinTransRNA achieves the highest accuracy and F1 among the compared methods on all three tasks (0.748/0.723, 0.914/0.908, and 0.819/0.823 for lncRNA, miRNA, and circRNA, respectively), surpassing RNALight and RNALoc-LM. Ablations show consistent gains over one-hot Transformer baselines and classical FCGR-based classifiers, while feature-space and curve-level analyses characterize the complementary contributions of the FCGR representation and the window-attention backbone. The main novelty lies in combining an order-traceable FCGR with hierarchical window attention under a reproducible protocol: five-seed repeats, confidence intervals, and paired tests quantify comparison uncertainty, and patch-embedding heatmaps with candidate motif screens expose the patterns prioritized by the model. The unified implementation, traceable sequence processing, and compact 4.25-million-parameter model provide a fixed-dimensional, auditable basis for RNA localization prediction and pattern prioritization.
Authors
- Jijun Tang (ORCID: https://orcid.org/0000-0002-6377-536X)
- Tzong-Yi Lee (ORCID: https://orcid.org/0009-0002-0283-7712)
- Leyi Wei (ORCID: https://orcid.org/0000-0003-1444-190X)
- Yixian Huang (ORCID: https://orcid.org/0009-0004-2601-2875)
- Lantian Yao (ORCID: https://orcid.org/0000-0003-4554-6827)
- Jiahui Guan (ORCID: https://orcid.org/0009-0001-0584-0637)
- Xingchen Liu (ORCID: https://orcid.org/0000-0002-5638-4053)
- Yen-Peng Chiu
- Peilin Xie (ORCID: https://orcid.org/0009-0000-6448-2908)
- Xi He (ORCID: https://orcid.org/0000-0003-3902-028X)
- Bowen Shi
- Zhihao Zhao
- Xiangrong Liu
Institutions
- National Yang Ming Chiao Tung University (TW)
- Xiamen University (CN)
- Chinese University of Hong Kong, Shenzhen (CN)
- Shenzhen Technology University (CN)
- Shenzhen University of Advanced Technology (CN)
- Macao Polytechnic University (MO)
- Xiamen University of Technology (CN)
- University of Hong Kong (HK)
Publication Details
- Journal
- Journal of Chemical Information and Modeling
- Published
- 2026-09-22
- DOI
- https://doi.org/10.1021/acs.jcim.6c02004
- Primary Topic
- Machine Learning in Bioinformatics
- Type
- article
- Field-Weighted Citation Impact
- 0.00