Development of a novel aggregated deep learning framework for small biological datasets using overlapping subsequences

Abstract The development of deep learning models and techniques to harness the potential of data mining in biological sequences with limited data availability is a vital need. This necessity arises from unique biological characteristics and technical constraints that limit access to sufficient quantities of high-quality data across many biological and genetic fields. Building on a data augmentation strategy that generates overlapping augmented subsequences, an aggregated CNN-LSTM-Attention-Residual architecture was developed for analysis at the original full-length sequence level. The model was independently tested and validated on three sets of regulatory sequences from evolutionarily diverse organisms, chloroplasts sharing a common ancestor, and prokaryotic sequences, comprising 100, 50, and 100 sequences per group, respectively, within each dataset. Trained on augmented subsequences, the model effectively aggregated and transferred high-level features back to the original full-length sequences. It achieved high performance with approximately 96% accuracy, recall, precision, F1-score, and AUC at the subsequence level. The model correctly distinguished the sequences of the distinct groups at the full-length sequence level across all datasets. The strong performance confirmed that an ensemble of features learned from short, overlapping subsequences provide an informative signature for classifying full-length regulatory regions. The combination of first-level training on augmented subsequences with dual-level validation—evaluating both subsequences and their corresponding full-length sequences— provides a framework for advance model development in limited-data biological sequence analysis.

Authors

Publication Details

Journal
Scientific Reports
Published
2026-09-16
DOI
https://doi.org/10.1038/s41598-026-69140-y
Primary Topic
Genomics and Phylogenetic Studies
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Development of a novel aggregated deep learning framework for small biological datasets using overlapping subsequences

Pär K. Ingvarsson, Naser Farrokhi, Mohammad Ali Abbasi-Vineh
Scientific Reports
Genomics and Phylogenetic Studies
article

Development of a novel aggregated deep learning framework for small biological datasets using overlapping subsequences

Pär K. Ingvarsson, Naser Farrokhi, Mohammad Ali Abbasi-Vineh
article en

Abstract

Abstract The development of deep learning models and techniques to harness the potential of data mining in biological sequences with limited data availability is a vital need. This necessity arises from unique biological characteristics and technical constraints that limit access to sufficient quantities of high-quality data across many biological and genetic fields. Building on a data augmentation strategy that generates overlapping augmented subsequences, an aggregated CNN-LSTM-Attention-Residual architecture was developed for analysis at the original full-length sequence level. The model was independently tested and validated on three sets of regulatory sequences from evolutionarily diverse organisms, chloroplasts sharing a common ancestor, and prokaryotic sequences, comprising 100, 50, and 100 sequences per group, respectively, within each dataset. Trained on augmented subsequences, the model effectively aggregated and transferred high-level features back to the original full-length sequences. It achieved high performance with approximately 96% accuracy, recall, precision, F1-score, and AUC at the subsequence level. The model correctly distinguished the sequences of the distinct groups at the full-length sequence level across all datasets. The strong performance confirmed that an ensemble of features learned from short, overlapping subsequences provide an informative signature for classifying full-length regulatory regions. The combination of first-level training on augmented subsequences with dual-level validation—evaluating both subsequences and their corresponding full-length sequences— provides a framework for advance model development in limited-data biological sequence analysis.

Scientific Reports
Openalex Percentile: Top 18%
Genomics and Phylogenetic Studies
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Development of a novel aggregated deep learning framework for small biological datasets using overlapping subsequences — Pär K. Ingvarsson, Naser Farrokhi, et al. · Scientific Reports (2026) | TGRS Research Map | TGRS