PolyAnalysis: a deep learning–based framework for nanopore-based poly(A) tail estimation and 3′-end regulation profiling
Abstract Background Polyadenylation at the 3′ end of mRNA regulates transcript stability, export, and translation, but poly(A) tail length and polyadenylation-site usage remain difficult to resolve from nanopore sequencing data because of homopolymer-associated errors and complex 3′ end architectures. We developed PolyAnalysis, a signal-level computational framework for estimating poly(A) tail length from raw Oxford Nanopore Technologies signals and linking per-read estimates to confidence-aware analyses of polyadenylation sites, alternative polyadenylation, and repetitive-element-associated 3′ ends. Results PolyAnalysis combines CNN–BiLSTM-based region classification with multi-task connectionist temporal classification decoding and adenine-aware training and decoding to infer poly(A)-proximal boundaries, tail length, and model-derived sequence features. Benchmarking against Nanopolish, tailfindr, and Dorado on controlled RNA002 and RNA004 datasets showed competitive tail-length accuracy and high callability. Component-wise ablation and bias-control analyses indicated that region classification, multi-task optimization, adenine-weighted loss, and A-rich decoding jointly contributed to performance. Application to human cell-line and vertebrate transcriptome datasets identified dataset-level variation in tail-length distributions, gene-level heterogeneity, and weak associations with steady-state transcript abundance. Confidence-based filtering separated database-matched, high-confidence novel-candidate and low-confidence polyadenylation sites, while ambiguity-aware endogenous-retrovirus analyses distinguished locus-level from family-level evidence. Conclusions PolyAnalysis provides a benchmarked framework for nanopore-based poly(A) tail-length estimation and a confidence-aware basis for exploratory 3′ end profiling. Tail-length estimation is the most extensively validated output, whereas sequence-composition, PAS/APA, and ERV-related results remain dependent on model assumptions, sequencing chemistry, and annotation quality and require further orthogonal validation.
Authors
- Qiuxiang Tian (ORCID: https://orcid.org/0009-0005-4410-5351)
- Bosheng Song (ORCID: https://orcid.org/0000-0002-1479-5399)
- Quan Zou
- Yansu Wang
Institutions
- University of Electronic Science and Technology of China (CN)
- Hunan University (CN)
- Hunan University of Finance and Economics (CN)
Publication Details
- Journal
- BMC Biology
- Published
- 2026-10-05
- DOI
- https://doi.org/10.1186/s12915-026-02749-7
- Primary Topic
- RNA Research and Splicing
- Type
- article
- Field-Weighted Citation Impact
- 0.00