The first reads are not a random sample: archive FASTQ files from aligned submissions can show head-of-file telomeric enrichment and inflate two-million-read estimates 7- to 10-fold

Whole-genome sequencing files are large, and a common shortcut is to analyse only their first few million reads. That is safe only if the order of reads in the file carries no information. We asked whether the shortcut distorts a rare signal, the share of reads made of telomeric repeats, in public human data from the European Nucleotide Archive (ENA). Each experiment's protocol and code were hashed before its outcomes were observed. Two complete-file benchmarks (EXP-001, EXP-002) compared the first 2,000,000 reads of each file with the complete file. In ten legacy files from five studies, EXP-001 did not support a practically large median excess error of the prefix over random samples of the same size; it did not establish that prefixes are safe. In EXP-002, a prespecified sample of 48 contemporary files, 4 failed badly, overestimating the telomeric signal 6.9- to 10.3-fold, and an independent implementation reproduced every checked count in nine of those files. A post hoc metadata analysis classified all four failures as aligned submissions (BAM, CRAM or an archive alignment record) rather than raw FASTQ: 4 of 5 aligned submissions failed versus 0 of 43 others. We then tested this marker prospectively on 48 new runs, fixing every prediction before download: telomeric reads were concentrated in the first million records, relative to the next million, in 15 of 24 BAM/CRAM submissions versus 1 of 24 FASTQ submissions. This experiment did not measure complete-file error. The profiles are consistent with coordinate order surviving the archive's conversion of a sorted alignment, which places reads from the chromosome 1 telomere first; the conversion step was not inspected. The practical rule: check the submission format before download, check read order before analysis, and use uniform sampling by default. We release a small tool for both checks; it reproduces the study's metadata labels for all 96 runs.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-03
DOI
https://doi.org/10.5281/zenodo.23123220
Primary Topic
Genomics and Phylogenetic Studies
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

The first reads are not a random sample: archive FASTQ files from aligned submissions can show head-of-file telomeric enrichment and inflate two-million-read estimates 7- to 10-fold

Mariusz Kulma
Zenodo (CERN European Organization for Nuclear Research)
Genomics and Phylogenetic Studies
preprint

The first reads are not a random sample: archive FASTQ files from aligned submissions can show head-of-file telomeric enrichment and inflate two-million-read estimates 7- to 10-fold

Mariusz Kulma
preprint en

Abstract

Whole-genome sequencing files are large, and a common shortcut is to analyse only their first few million reads. That is safe only if the order of reads in the file carries no information. We asked whether the shortcut distorts a rare signal, the share of reads made of telomeric repeats, in public human data from the European Nucleotide Archive (ENA). Each experiment's protocol and code were hashed before its outcomes were observed. Two complete-file benchmarks (EXP-001, EXP-002) compared the first 2,000,000 reads of each file with the complete file. In ten legacy files from five studies, EXP-001 did not support a practically large median excess error of the prefix over random samples of the same size; it did not establish that prefixes are safe. In EXP-002, a prespecified sample of 48 contemporary files, 4 failed badly, overestimating the telomeric signal 6.9- to 10.3-fold, and an independent implementation reproduced every checked count in nine of those files. A post hoc metadata analysis classified all four failures as aligned submissions (BAM, CRAM or an archive alignment record) rather than raw FASTQ: 4 of 5 aligned submissions failed versus 0 of 43 others. We then tested this marker prospectively on 48 new runs, fixing every prediction before download: telomeric reads were concentrated in the first million records, relative to the next million, in 15 of 24 BAM/CRAM submissions versus 1 of 24 FASTQ submissions. This experiment did not measure complete-file error. The profiles are consistent with coordinate order surviving the archive's conversion of a sorted alignment, which places reads from the chromosome 1 telomere first; the conversion step was not inspected. The practical rule: check the submission format before download, check read order before analysis, and use uniform sampling by default. We release a small tool for both checks; it reproduces the study's metadata labels for all 96 runs.

Zenodo (CERN European Organization for Nuclear Research)
Genomics and Phylogenetic Studies
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

The first reads are not a random sample: archive FASTQ files from aligned submissions can show head-of-file telomeric enrichment and inflate two-million-read estimates 7- to 10-fold — Mariusz Kulma · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS