The first reads are not a random sample: archive FASTQ files from aligned submissions can show head-of-file telomeric enrichment and inflate two-million-read estimates 7- to 10-fold
Whole-genome sequencing files are large, and a common shortcut is to analyse only their first few million reads. That is safe only if the order of reads in the file carries no information. We asked whether the shortcut distorts a rare signal, the share of reads made of telomeric repeats, in public human data from the European Nucleotide Archive (ENA). Each experiment's protocol and code were hashed before its outcomes were observed. Two complete-file benchmarks (EXP-001, EXP-002) compared the first 2,000,000 reads of each file with the complete file. In ten legacy files from five studies, EXP-001 did not support a practically large median excess error of the prefix over random samples of the same size; it did not establish that prefixes are safe. In EXP-002, a prespecified sample of 48 contemporary files, 4 failed badly, overestimating the telomeric signal 6.9- to 10.3-fold, and an independent implementation reproduced every checked count in nine of those files. A post hoc metadata analysis classified all four failures as aligned submissions (BAM, CRAM or an archive alignment record) rather than raw FASTQ: 4 of 5 aligned submissions failed versus 0 of 43 others. We then tested this marker prospectively on 48 new runs, fixing every prediction before download: telomeric reads were concentrated in the first million records, relative to the next million, in 15 of 24 BAM/CRAM submissions versus 1 of 24 FASTQ submissions. This experiment did not measure complete-file error. The profiles are consistent with coordinate order surviving the archive's conversion of a sorted alignment, which places reads from the chromosome 1 telomere first; the conversion step was not inspected. The practical rule: check the submission format before download, check read order before analysis, and use uniform sampling by default. We release a small tool for both checks; it reproduces the study's metadata labels for all 96 runs.
Authors
- Mariusz Kulma (ORCID: https://orcid.org/0009-0000-5550-8723)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-03
- DOI
- https://doi.org/10.5281/zenodo.23123220
- Primary Topic
- Genomics and Phylogenetic Studies
- Type
- preprint