From public genome data to biological insight: A computational workflow for in silico restriction site analyses

This poster was presented on the TRA Workshop "Great Data, Great Science" at the University of Bonn on 06 October 2026. --- In computational biology, the availability and careful handling of high-quality digital data are essential for obtaining reliable and biologically meaningful results. In silico analyses depend on accurate sequence information, and many investigations begin with publicly available genome assemblies retrieved from resources such as the National Center for Biotechnology Information (NCBI; https://www.ncbi.nlm.nih.gov) or Ensembl (https://ensemblgenomes.org). However, biological questions often require specialized analytical strategies that are not directly supported by standard databases or existing software. Consequently, custom computational workflows - implemented, for example, in Python or R - are frequently necessary to process genomic data and extract relevant biological patterns. Our poster presents such a workflow for the in silico analysis of restriction sites in complete genome sequences. The workflow begins with the retrieval and preparation of published genomic data, continues with the development and application of custom scripts, and culminates in the interpretation of the resulting biological patterns. In addition, we describe the subsequent publication and documentation of both the analysis code and the generated data, thereby supporting transparency, reproducibility, and reuse. The biological objective of this study is to investigate the distribution of restriction fragments generated by different restriction enzymes and to compare these distributions with those expected from random sequences. Biological sequences are not randomly assembled: their composition and organization are shaped by evolutionary forces such as selection, mutation, recombination, and sequence duplication. Therefore, deviations in restriction-fragment distributions may provide an indirect indication of underlying genomic organization and non-random sequence structure. Our computational analyses demonstrate that, for most combinations of restriction enzyme and genome sequence examined, the number of fragments per megabase differs by more than (10%) between biological and corresponding random sequences. These deviations occur in both directions: substantially increased values are approximately as common as substantially decreased values. We did not observe a consistent species-specific or restriction-enzyme-specific effect. Nevertheless, the results reveal a clear influence of GC content, both at the level of the restriction site and across the analyzed genome sequence. Furthermore, unlike the random controls, the biological genome sequences display distinct peaks in their fragment-length distributions. These recurring peaks may reflect repetitive genomic elements, including transposable elements and other duplicated sequences.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-06
DOI
https://doi.org/10.5281/zenodo.23104295
Primary Topic
Genomics and Phylogenetic Studies
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

From public genome data to biological insight: A computational workflow for in silico restriction site analyses

Heiko Schoof, Lucia Vedder
Zenodo (CERN European Organization for Nuclear Research)
Genomics and Phylogenetic Studies
article

From public genome data to biological insight: A computational workflow for in silico restriction site analyses

Heiko Schoof, Lucia Vedder
article en

Abstract

This poster was presented on the TRA Workshop "Great Data, Great Science" at the University of Bonn on 06 October 2026. --- In computational biology, the availability and careful handling of high-quality digital data are essential for obtaining reliable and biologically meaningful results. In silico analyses depend on accurate sequence information, and many investigations begin with publicly available genome assemblies retrieved from resources such as the National Center for Biotechnology Information (NCBI; https://www.ncbi.nlm.nih.gov) or Ensembl (https://ensemblgenomes.org). However, biological questions often require specialized analytical strategies that are not directly supported by standard databases or existing software. Consequently, custom computational workflows - implemented, for example, in Python or R - are frequently necessary to process genomic data and extract relevant biological patterns. Our poster presents such a workflow for the in silico analysis of restriction sites in complete genome sequences. The workflow begins with the retrieval and preparation of published genomic data, continues with the development and application of custom scripts, and culminates in the interpretation of the resulting biological patterns. In addition, we describe the subsequent publication and documentation of both the analysis code and the generated data, thereby supporting transparency, reproducibility, and reuse. The biological objective of this study is to investigate the distribution of restriction fragments generated by different restriction enzymes and to compare these distributions with those expected from random sequences. Biological sequences are not randomly assembled: their composition and organization are shaped by evolutionary forces such as selection, mutation, recombination, and sequence duplication. Therefore, deviations in restriction-fragment distributions may provide an indirect indication of underlying genomic organization and non-random sequence structure. Our computational analyses demonstrate that, for most combinations of restriction enzyme and genome sequence examined, the number of fragments per megabase differs by more than (10%) between biological and corresponding random sequences. These deviations occur in both directions: substantially increased values are approximately as common as substantially decreased values. We did not observe a consistent species-specific or restriction-enzyme-specific effect. Nevertheless, the results reveal a clear influence of GC content, both at the level of the restriction site and across the analyzed genome sequence. Furthermore, unlike the random controls, the biological genome sequences display distinct peaks in their fragment-length distributions. These recurring peaks may reflect repetitive genomic elements, including transposable elements and other duplicated sequences.

Zenodo (CERN European Organization for Nuclear Research)
University of Bonn (DE)
Openalex Percentile: Top 21%
Genomics and Phylogenetic Studies
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.