Improving RNA Secondary Structure Prediction Through Expanded Training Data

In recent years, deep learning has revolutionized protein structure prediction, achieving remarkable speed and accuracy. RNA structure prediction, however, has lagged behind. Although several methods have shown some success in predicting RNA secondary and tertiary structures, none have reached the accuracy observed with contemporary protein models. The lack of success of these RNA structure prediction models has been proposed to be due to limited high-quality structural information that can be used as training data. To probe this proposed limitation, we developed a large and diverse dataset comprising paired RNA sequences and their corresponding secondary structures. We assess the utility of this enhanced dataset by retraining on a deep learning model, SincFold. We find that SincFold exhibited improved generalization to some previously unseen RNA families, enhancing its capability to predict accurate de novo RNA secondary structures. The RNASSTR dataset provides a substantial advance for RNA structure modeling, laying a strong foundation for the development of future RNA secondary structure prediction algorithms.

Authors

Institutions

Publication Details

Journal
RNA
Published
2026-09-01
DOI
https://doi.org/10.1261/rna.081259.126
Citations
2
Primary Topic
RNA and protein synthesis mechanisms
Type
article
Field-Weighted Citation Impact
5.88

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Improving RNA Secondary Structure Prediction Through Expanded Training Data

Agni Rajinikanth, J.H.D. Cate, C. A. Meredith, Conner J. Langeberg et al.
2 citations
RNA
RNA and protein synthesis mechanisms
5.88
article

Improving RNA Secondary Structure Prediction Through Expanded Training Data

Agni Rajinikanth, J.H.D. Cate, C. A. Meredith, Conner J. Langeberg, Jennifer A. Doudna, Taehan Kim, Dimple Amitha Garuadapuri, Roma Nagle
article en
2 citations

Abstract

In recent years, deep learning has revolutionized protein structure prediction, achieving remarkable speed and accuracy. RNA structure prediction, however, has lagged behind. Although several methods have shown some success in predicting RNA secondary and tertiary structures, none have reached the accuracy observed with contemporary protein models. The lack of success of these RNA structure prediction models has been proposed to be due to limited high-quality structural information that can be used as training data. To probe this proposed limitation, we developed a large and diverse dataset comprising paired RNA sequences and their corresponding secondary structures. We assess the utility of this enhanced dataset by retraining on a deep learning model, SincFold. We find that SincFold exhibited improved generalization to some previously unseen RNA families, enhancing its capability to predict accurate de novo RNA secondary structures. The RNASSTR dataset provides a substantial advance for RNA structure modeling, laying a strong foundation for the development of future RNA secondary structure prediction algorithms.

RNA
QB3 (US), Howard Hughes Medical Institute (US), Lawrence Berkeley National Laboratory (US), Innovative Genomics Institute (US), University of California, Berkeley (US)
National Science Foundation, Howard Hughes Medical Institute, U.S. Department of Energy, Innovative Genomics Institute, Emerson Collective, National Institutes of Health, National Institute of Allergy and Infectious Diseases, National Institute of Neurological Disorders and Stroke, Lawrence Livermore National Laboratory
Openalex Percentile: Top 7%
RNA and protein synthesis mechanisms
5.88
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.