Distributed, Not Sparse: Decoder-Basis Geometry of Toxin-Related Directions in a Protein Language Model
Sparse autoencoders (SAEs) are increasingly used to identify interpretable features in protein language models, but a failure to recover stable sparse features can have at least two explanations: the relevant biological information may be absent from the underlying representation, or it may be present but poorly aligned with the learned sparse basis. This study distinguishes these possibilities for toxin-related information in ESM-2. Prior experiments found strong linear accessibility of toxin-related information in raw ESM-2 representations, with median held-out AUROC reaching 0.965 at layer 24 compared with 0.587 for a sequence-length and amino-acid-composition baseline, while an InterPLM SAE at layer 18 failed preregistered stability criteria. We therefore examined the geometry of supervised toxin-related directions relative to the frozen InterPLM layer-18 decoder. Using orthogonal matching pursuit across biological, canonical label-null, C-matched null, and isotropic reference directions, 32 decoder atoms reconstructed a median 25.68% of the biological direction, compared with 24.33% for C-matched null directions and 24.12% for isotropic directions. The biological decoder-weight distribution had an effective support of approximately 6,374 of 10,231 usable decoder atoms. All biological directions required more than 64 atoms to reach 50% reconstruction, more than 128 to reach 80%, and more than 256 to reach 90%. The dominant result is therefore strong linear accessibility without compact SAE-basis decomposition. The earlier SAE stability failure cannot be attributed to absence of linearly accessible toxin-related information; instead, the distributed decoder geometry provides a plausible geometric explanation for why a compact recurrent feature set was not recovered. Modest preferential decoder alignment relative to controls is visible in point estimates, but biological bootstrap intervals primarily characterize solver-seed variation conditional on a fixed dataset and should not be interpreted as biological sampling uncertainty. These findings identify a contrasting regime to recent positive protein-SAE results and motivate further evaluation of when biological functions are sparsely localized versus distributed in protein foundation-model representations. The study is descriptive and discovery-only and does not establish cross-family generalization, causal biological mechanisms, or operational hazardous-sequence screening performance.
Authors
- Allan Ochola (ORCID: https://orcid.org/0000-0002-0480-0123)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-09
- DOI
- https://doi.org/10.5281/zenodo.22681898
- Primary Topic
- Machine Learning in Bioinformatics
- Type
- preprint