k -mer-based Upstream Preprocessing of long reads for Isoform Discovery

Eukaryotic genes can encode multiple protein isoforms based on alternative splicing of their transcribed regions. Most modern novel isoform discovery methods function by identifying and assembling exon splice junctions from an RNA-seq sample. However, splice junctions can only be accurately annotated with time-intensive dynamic programming alignment. This manuscript introduces KuPID, a method for preprocessing long RNA-seq reads with the goal of better identifying novel isoform transcripts. KuPID utilizes k -mer sketching as a prefilter to quickly pseudo-align reads to known reference isoforms. Full alignment need only then be applied to reads that are most relevant to isoform discovery. Not only does KuPID speed up the discovery pipeline, it also increases downstream accuracy by filtering out extraneous reads. KuPID preprocessing simultaneously increases the f1 accuracy of isoform discovery pipelines by up to 11.6 points while decreasing the runtime by a factor of 2-3×;. An optional mode permits a KuPID sample to be paired with both isoform discovery and transcript quantification.

Authors

Institutions

Publication Details

Journal
Genome Research
Published
2026-09-14
DOI
https://doi.org/10.1101/gr.282250.126
Primary Topic
RNA Research and Splicing
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

k -mer-based Upstream Preprocessing of long reads for Isoform Discovery

Yun William Yu, Molly Borowiak
Genome Research
RNA Research and Splicing
preprint

k -mer-based Upstream Preprocessing of long reads for Isoform Discovery

Yun William Yu, Molly Borowiak
preprint en

Abstract

Eukaryotic genes can encode multiple protein isoforms based on alternative splicing of their transcribed regions. Most modern novel isoform discovery methods function by identifying and assembling exon splice junctions from an RNA-seq sample. However, splice junctions can only be accurately annotated with time-intensive dynamic programming alignment. This manuscript introduces KuPID, a method for preprocessing long RNA-seq reads with the goal of better identifying novel isoform transcripts. KuPID utilizes k -mer sketching as a prefilter to quickly pseudo-align reads to known reference isoforms. Full alignment need only then be applied to reads that are most relevant to isoform discovery. Not only does KuPID speed up the discovery pipeline, it also increases downstream accuracy by filtering out extraneous reads. KuPID preprocessing simultaneously increases the f1 accuracy of isoform discovery pipelines by up to 11.6 points while decreasing the runtime by a factor of 2-3×;. An optional mode permits a KuPID sample to be paired with both isoform discovery and transcript quantification.

Genome Research
Carnegie Mellon University (US)
RNA Research and Splicing
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

k -mer-based Upstream Preprocessing of long reads for Isoform Discovery — Yun William Yu, Molly Borowiak · Genome Research (2026) | TGRS Research Map | TGRS