What a local encoder cannot see: spectral fibres of the human proteome and the representation of perturbagens
UniPert-G2CP (Li et al., Cell 189, 2026) represents genetic perturbagens by their protein sequences and chemical perturbagens by ECFP4 fingerprints, and shows that its protein encoder outperforms composition-type encoders such as pseudo amino acid composition (PseAAC). We examine this comparison through algebraic genomics, the combinatorics of sequences with prescribed local content. An encoder is local of order k if it depends on a sequence only through its multiset of k-mers. Composition, k-mer composition, PseAAC with lag λ ≤ k−1 and every mean-pooled convolutional encoder of window k are local; so is ECFP, in the molecular-graph sense. A local encoder is constant on the spectral fibre of a sequence, the set of sequences with the same k-mer multiset. The size of the fibre is an Eulerian-trail count, which we compute exactly by the BEST theorem after contracting forced arcs, for all 20,431 reviewed human proteins. Separation is not resolution. No two distinct human proteins share a k-mer spectrum for any k ≥ 2, so every local encoder of order at least 2 separates the reference proteome. Yet the median protein shares its 3-mer spectrum with 2^239 other sequences and its 4-mer spectrum with 2^19. The spectral resolution κ, the least k beyond which the spectrum determines the sequence, has median 5 and exceeds 10 for 854 proteins. The unresolved order lives in duplication-driven families. The largest κ occur in the hominoid-specific segmental-duplication families NBPF (median κ = 46, up to 951), GOLGA6, NPIP and POTE, and in keratin-associated proteins, collagens and mucins. KRAB zinc-finger proteins are 3% of the proteome but 41% of the proteins with κ > 10. Explicit collisions. For 12 human proteins we construct a different sequence, altered at up to 15% of its positions, with the same 25-spectrum; it receives the identical 68-dimensional PseAAC vector (λ = 24, the setting of the UniPert benchmark) and the identical output of a window-25 convolutional encoder. On the DNA side, a strand-agnostic local encoder is constant on the set of double-stranded reconstructions, which for bacteriophage φX174 at k = 10 has 31,925,753,246,212 elements; computing its size is #P-hard in general. The record contains the paper, the code, the UniProt snapshot used, the per-protein results and the figures. All computations are exact except where stated, and the verification script reproduces every exact number in the paper in about five minutes on 16 cores.
Authors
- Zhengyi Chen
- Ruqing Chen
Institutions
- Guilin Medical University (CN)
- Energoservis (Czechia) (CZ)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-25
- DOI
- https://doi.org/10.5281/zenodo.22950094
- Primary Topic
- Bioinformatics and Genomic Networks
- Type
- preprint