The first reads are not a random sample: archive FASTQ files from aligned submissions can inflate two-million-read telomeric estimates 7- to 10-fold
Whole-genome sequencing files are large, and a common shortcut is to analyse only their first few million reads. That is safe only if the order of reads in the file carries no information. We asked whether the shortcut distorts a rare signal, the share of reads made of telomeric repeats, in public human data from the European Nucleotide Archive. Each experiment's protocol and code were fixed, and their cryptographic hashes recorded, before its outcomes were observed. In brief: in our sample of files from contemporary instruments, the shortcut failed badly in a minority of files. Every failing file had been submitted to the archive as reads already aligned to a reference genome, not as raw reads, and this can be seen in the archive's metadata record before download. In ten older files, we could not show that the first two million reads were typically much worse than random samples of the same size; that does not make them safe. In a prespecified sample of 48 files from contemporary instruments, the same shortcut overestimated the telomeric signal 7- to 10-fold in 4 files; their link to aligned submissions (BAM, CRAM or an archive alignment record) was found in a later analysis, not planned in advance. We then tested this marker on new runs, with every prediction recorded before download: telomeric reads were concentrated at the start of the file in 15 of 24 aligned submissions and in 1 of 24 raw-read (FASTQ) submissions. This last test did not measure the size of the error. The read-order pattern is consistent with the archive keeping the genome-position order of a sorted alignment when it converted the alignment to FASTQ, which places reads from the chromosome 1 telomere first; we did not inspect the conversion step. The practical rule: check the submission format before download, check read order before analysis, and by default sample reads uniformly at random across the whole file. We release a small tool for both checks. Version 2 (2026-10-06): the title, abstract, first paragraph of the Introduction and the Conclusion were rewritten for readability. Methods, results, tables, figures, data and code are unchanged from version 1.
Authors
- Mariusz Kulma (ORCID: https://orcid.org/0009-0000-5550-8723)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-06
- DOI
- https://doi.org/10.5281/zenodo.23188598
- Primary Topic
- Genomics and Phylogenetic Studies
- Type
- preprint