When does sequence order matter? An information-theoretic taxonomy of biological representation tasks and the limits of local encoding

Biological sequence models range from amino acid composition and k-mer counters to convolutional networks and billion-parameter language models. When is a local representation sufficient, and when is global order indispensable? This closing paper of an eight-part series answers the question with an exact information-theoretic taxonomy across the 20,431 reviewed human proteins. For any sequence S of length L over an alphabet of size q and any local order k, the specification cost L log2 q splits into three non-negative telescoping terms: the composition constraint L log2 q − B1(S), the resolved local-syntax information B1(S) − Bk(S), and the unresolved global-order entropy Bk(S) = log2 |Fk(S)|, the fibre entropy of the k-spectrum. Every local encoder of order k, from k-mer composition and pseudo amino acid composition to mean-pooled convolutional networks of window k, is provably blind to the third term. Evaluated proteome-wide for 1 ≤ k ≤ 50, composition fixes a median of 0.3947 bits per residue. Tripeptides resolve 83.34% of the permutation entropy B1(S) on average and pentapeptides 99.75%, yet 99.64% of proteins retain a non-trivial fibre at k = 3, 46.52% at k = 5 and 4.18% (854 proteins) at k = 10. Six functional classes fall into three regimes. Catalytic enzymes (n = 3,449), receptors, channels and transporters (n = 1,816) and other globular proteins (n = 14,153) have median spectral resolution κ = 5 and 95th percentile κ ≤ 8, following the random-sequence law κ ≈ 2 log20 L. The unresolved tail at k > 10 is concentrated in C2H2 zinc-finger regulators (n = 747, 49.13% with κ > 10), structural tandem-repeat proteins (n = 199, 39.70%) and hominoid segmental-duplication arrays (n = 67, 74.63%, 95th percentile κ = 257). Combining the decomposition with the companion results on convolutional receptive fields, ClinVar copy ambiguity, C2H2 zinc-finger reorderings, FibreBench drift ratios and reverse-complement and fingerprint symmetry quotients yields a four-step decision rule for choosing sequence encoders: check the spectral resolution of the target class, test phenotypic sufficiency on the fibre, choose an escaping mechanism that matches the read-out, and avoid global mean pooling when domain order or copy index is the signal. This record contains the paper (PDF and LaTeX source), the code (Python), the per-protein fibre catalogue it reads, the results and the figure. Companion papers: doi:10.5281/zenodo.22950095, doi:10.5281/zenodo.22959211, doi:10.5281/zenodo.22960439, doi:10.5281/zenodo.22963882, doi:10.5281/zenodo.22966020, doi:10.5281/zenodo.22998488, doi:10.5281/zenodo.23001957, doi:10.5281/zenodo.23002923. Monograph: doi:10.5281/zenodo.22943487. Code: https://github.com/Ruqing1963/order-taxonomy

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-28
DOI
https://doi.org/10.5281/zenodo.23003752
Primary Topic
Machine Learning in Bioinformatics
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

When does sequence order matter? An information-theoretic taxonomy of biological representation tasks and the limits of local encoding

Zhengyi Chen, Ruqing Chen
Zenodo (CERN European Organization for Nuclear Research)
Machine Learning in Bioinformatics
preprint

When does sequence order matter? An information-theoretic taxonomy of biological representation tasks and the limits of local encoding

Zhengyi Chen, Ruqing Chen
preprint en

Abstract

Biological sequence models range from amino acid composition and k-mer counters to convolutional networks and billion-parameter language models. When is a local representation sufficient, and when is global order indispensable? This closing paper of an eight-part series answers the question with an exact information-theoretic taxonomy across the 20,431 reviewed human proteins. For any sequence S of length L over an alphabet of size q and any local order k, the specification cost L log2 q splits into three non-negative telescoping terms: the composition constraint L log2 q − B1(S), the resolved local-syntax information B1(S) − Bk(S), and the unresolved global-order entropy Bk(S) = log2 |Fk(S)|, the fibre entropy of the k-spectrum. Every local encoder of order k, from k-mer composition and pseudo amino acid composition to mean-pooled convolutional networks of window k, is provably blind to the third term. Evaluated proteome-wide for 1 ≤ k ≤ 50, composition fixes a median of 0.3947 bits per residue. Tripeptides resolve 83.34% of the permutation entropy B1(S) on average and pentapeptides 99.75%, yet 99.64% of proteins retain a non-trivial fibre at k = 3, 46.52% at k = 5 and 4.18% (854 proteins) at k = 10. Six functional classes fall into three regimes. Catalytic enzymes (n = 3,449), receptors, channels and transporters (n = 1,816) and other globular proteins (n = 14,153) have median spectral resolution κ = 5 and 95th percentile κ ≤ 8, following the random-sequence law κ ≈ 2 log20 L. The unresolved tail at k > 10 is concentrated in C2H2 zinc-finger regulators (n = 747, 49.13% with κ > 10), structural tandem-repeat proteins (n = 199, 39.70%) and hominoid segmental-duplication arrays (n = 67, 74.63%, 95th percentile κ = 257). Combining the decomposition with the companion results on convolutional receptive fields, ClinVar copy ambiguity, C2H2 zinc-finger reorderings, FibreBench drift ratios and reverse-complement and fingerprint symmetry quotients yields a four-step decision rule for choosing sequence encoders: check the spectral resolution of the target class, test phenotypic sufficiency on the fibre, choose an escaping mechanism that matches the read-out, and avoid global mean pooling when domain order or copy index is the signal. This record contains the paper (PDF and LaTeX source), the code (Python), the per-protein fibre catalogue it reads, the results and the figure. Companion papers: doi:10.5281/zenodo.22950095, doi:10.5281/zenodo.22959211, doi:10.5281/zenodo.22960439, doi:10.5281/zenodo.22963882, doi:10.5281/zenodo.22966020, doi:10.5281/zenodo.22998488, doi:10.5281/zenodo.23001957, doi:10.5281/zenodo.23002923. Monograph: doi:10.5281/zenodo.22943487. Code: https://github.com/Ruqing1963/order-taxonomy

Zenodo (CERN European Organization for Nuclear Research)
Guilin Medical University (CN)
Machine Learning in Bioinformatics
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.