When does sequence order matter? An information-theoretic taxonomy of biological representation tasks and the limits of local encoding
Biological sequence models range from amino acid composition and k-mer counters to convolutional networks and billion-parameter language models. When is a local representation sufficient, and when is global order indispensable? This closing paper of an eight-part series answers the question with an exact information-theoretic taxonomy across the 20,431 reviewed human proteins. For any sequence S of length L over an alphabet of size q and any local order k, the specification cost L log2 q splits into three non-negative telescoping terms: the composition constraint L log2 q − B1(S), the resolved local-syntax information B1(S) − Bk(S), and the unresolved global-order entropy Bk(S) = log2 |Fk(S)|, the fibre entropy of the k-spectrum. Every local encoder of order k, from k-mer composition and pseudo amino acid composition to mean-pooled convolutional networks of window k, is provably blind to the third term. Evaluated proteome-wide for 1 ≤ k ≤ 50, composition fixes a median of 0.3947 bits per residue. Tripeptides resolve 83.34% of the permutation entropy B1(S) on average and pentapeptides 99.75%, yet 99.64% of proteins retain a non-trivial fibre at k = 3, 46.52% at k = 5 and 4.18% (854 proteins) at k = 10. Six functional classes fall into three regimes. Catalytic enzymes (n = 3,449), receptors, channels and transporters (n = 1,816) and other globular proteins (n = 14,153) have median spectral resolution κ = 5 and 95th percentile κ ≤ 8, following the random-sequence law κ ≈ 2 log20 L. The unresolved tail at k > 10 is concentrated in C2H2 zinc-finger regulators (n = 747, 49.13% with κ > 10), structural tandem-repeat proteins (n = 199, 39.70%) and hominoid segmental-duplication arrays (n = 67, 74.63%, 95th percentile κ = 257). Combining the decomposition with the companion results on convolutional receptive fields, ClinVar copy ambiguity, C2H2 zinc-finger reorderings, FibreBench drift ratios and reverse-complement and fingerprint symmetry quotients yields a four-step decision rule for choosing sequence encoders: check the spectral resolution of the target class, test phenotypic sufficiency on the fibre, choose an escaping mechanism that matches the read-out, and avoid global mean pooling when domain order or copy index is the signal. This record contains the paper (PDF and LaTeX source), the code (Python), the per-protein fibre catalogue it reads, the results and the figure. Companion papers: doi:10.5281/zenodo.22950095, doi:10.5281/zenodo.22959211, doi:10.5281/zenodo.22960439, doi:10.5281/zenodo.22963882, doi:10.5281/zenodo.22966020, doi:10.5281/zenodo.22998488, doi:10.5281/zenodo.23001957, doi:10.5281/zenodo.23002923. Monograph: doi:10.5281/zenodo.22943487. Code: https://github.com/Ruqing1963/order-taxonomy
Authors
- Zhengyi Chen
- Ruqing Chen
Institutions
- Guilin Medical University (CN)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-28
- DOI
- https://doi.org/10.5281/zenodo.23003751
- Primary Topic
- Machine Learning in Bioinformatics
- Type
- preprint