Separation without resolution: the spectral fibre geometry of the human proteome and the blind spots of local sequence encoders
Protein representations are judged by how well they tell known proteins apart. We show that this criterion cannot detect what an encoder of a common form is unable to see. An encoder is local of order k if it depends on a sequence only through its multiset of k-mers. Amino acid and k-mer composition, pseudo amino acid composition with lag below k, and every mean-pooled convolutional network of receptive field k are local. Such an encoder factors through the quotient of sequence space by spectral equivalence, and it is constant on the spectral fibre of every sequence. We compute fibre sizes exactly for all 20,431 reviewed human proteins, as Eulerian-trail counts in de Bruijn multigraphs by the BEST theorem after contracting forced arcs. Separation and resolution come apart sharply. No two distinct human proteins share a k-mer spectrum for any k ≥ 2, so every local encoder of order two or more separates the reference proteome perfectly. Yet the median protein shares its 3-mer spectrum with about 2^239 other sequences and its 4-mer spectrum with about 2^19; only 0.4% of proteins are determined by their 3-mer spectrum. The spectral resolution κ, the order beyond which the spectrum determines the sequence, has median 5 and exceeds 10 for 854 proteins. These blind spots do not follow length. They sit in repeat architectures. Hominoid segmental-duplication families (NBPF, GOLGA6, NPIP, POTE) make up 0.3% of the proteome but 5.9% of the proteins with κ > 10, reaching κ = 951 in NBPF20. C2H2 zinc-finger proteins make up 3.0% and 40.7%. Five repeat architectures together account for 4% of the proteome and 56% of that tail. For 12 human proteins, an explicit different sequence with the same 25-spectrum receives the identical 68-dimensional PseAAC vector and the identical output of a window-25 convolutional encoder. Benchmarks on a reference set measure separation. Resolution has to be measured on fibres. The statements concern local encoders; transformer language models such as ESM-2 are not local, and no trained model was evaluated. This record contains the paper (PDF and LaTeX source), the code (Python), the UniProt snapshot used, the per-protein results and the figures. The computations are shared with the companion report doi:10.5281/zenodo.22950095, which examines one perturbation-modelling framework; the present paper makes the separation–resolution gap its subject and adds the atlas by repeat architecture. Code: https://github.com/Ruqing1963/separation-resolution
Authors
- Zhengyi Chen
- Ruqing Chen
Institutions
- Guilin Medical University (CN)
- Energoservis (Czechia) (CZ)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-25
- DOI
- https://doi.org/10.5281/zenodo.22959210
- Primary Topic
- Machine Learning in Bioinformatics
- Type
- preprint