The first reads are not a random sample: archive FASTQ files from aligned submissions can inflate two-million-read telomeric estimates 7- to 10-fold

Whole-genome sequencing files are large, and a common shortcut is to analyse only their first few million reads. That is safe only if the order of reads in the file carries no information. We asked whether the shortcut distorts a rare signal, the share of reads made of telomeric repeats, in public human data from the European Nucleotide Archive. Each experiment's protocol and code were fixed, and their cryptographic hashes recorded, before its outcomes were observed. In brief: in our sample of files from contemporary instruments, the shortcut failed badly in a minority of files. Every failing file had been submitted to the archive as reads already aligned to a reference genome, not as raw reads, and this can be seen in the archive's metadata record before download. In ten older files, we could not show that the first two million reads were typically much worse than random samples of the same size; that does not make them safe. In a prespecified sample of 48 files from contemporary instruments, the same shortcut overestimated the telomeric signal 7- to 10-fold in 4 files; their link to aligned submissions (BAM, CRAM or an archive alignment record) was found in a later analysis, not planned in advance. We then tested this marker on new runs, with every prediction recorded before download: telomeric reads were concentrated at the start of the file in 15 of 24 aligned submissions and in 1 of 24 raw-read (FASTQ) submissions. This last test did not measure the size of the error. The read-order pattern is consistent with the archive keeping the genome-position order of a sorted alignment when it converted the alignment to FASTQ, which places reads from the chromosome 1 telomere first; we did not inspect the conversion step. The practical rule: check the submission format before download, check read order before analysis, and by default sample reads uniformly at random across the whole file. We release a small tool for both checks. Version 2 (2026-10-06): the title, abstract, first paragraph of the Introduction and the Conclusion were rewritten for readability. Methods, results, tables, figures, data and code are unchanged from version 1.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-06
DOI
https://doi.org/10.5281/zenodo.23188598
Primary Topic
Genomics and Phylogenetic Studies
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

The first reads are not a random sample: archive FASTQ files from aligned submissions can inflate two-million-read telomeric estimates 7- to 10-fold

Mariusz Kulma
Zenodo (CERN European Organization for Nuclear Research)
Genomics and Phylogenetic Studies
preprint

The first reads are not a random sample: archive FASTQ files from aligned submissions can inflate two-million-read telomeric estimates 7- to 10-fold

Mariusz Kulma
preprint en

Abstract

Whole-genome sequencing files are large, and a common shortcut is to analyse only their first few million reads. That is safe only if the order of reads in the file carries no information. We asked whether the shortcut distorts a rare signal, the share of reads made of telomeric repeats, in public human data from the European Nucleotide Archive. Each experiment's protocol and code were fixed, and their cryptographic hashes recorded, before its outcomes were observed. In brief: in our sample of files from contemporary instruments, the shortcut failed badly in a minority of files. Every failing file had been submitted to the archive as reads already aligned to a reference genome, not as raw reads, and this can be seen in the archive's metadata record before download. In ten older files, we could not show that the first two million reads were typically much worse than random samples of the same size; that does not make them safe. In a prespecified sample of 48 files from contemporary instruments, the same shortcut overestimated the telomeric signal 7- to 10-fold in 4 files; their link to aligned submissions (BAM, CRAM or an archive alignment record) was found in a later analysis, not planned in advance. We then tested this marker on new runs, with every prediction recorded before download: telomeric reads were concentrated at the start of the file in 15 of 24 aligned submissions and in 1 of 24 raw-read (FASTQ) submissions. This last test did not measure the size of the error. The read-order pattern is consistent with the archive keeping the genome-position order of a sorted alignment when it converted the alignment to FASTQ, which places reads from the chromosome 1 telomere first; we did not inspect the conversion step. The practical rule: check the submission format before download, check read order before analysis, and by default sample reads uniformly at random across the whole file. We release a small tool for both checks. Version 2 (2026-10-06): the title, abstract, first paragraph of the Introduction and the Conclusion were rewritten for readability. Methods, results, tables, figures, data and code are unchanged from version 1.

Zenodo (CERN European Organization for Nuclear Research)
Genomics and Phylogenetic Studies
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

The first reads are not a random sample: archive FASTQ files from aligned submissions can inflate two-million-read telomeric estimates 7- to 10-fold — Mariusz Kulma · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS