Distributed, Not Sparse: Decoder-Basis Geometry of Toxin-Related Directions in a Protein Language Model

Sparse autoencoders (SAEs) are increasingly used to identify interpretable features in protein language models, but a failure to recover stable sparse features can have at least two explanations: the relevant biological information may be absent from the underlying representation, or it may be present but poorly aligned with the learned sparse basis. This study distinguishes these possibilities for toxin-related information in ESM-2. Prior experiments found strong linear accessibility of toxin-related information in raw ESM-2 representations, with median held-out AUROC reaching 0.965 at layer 24 compared with 0.587 for a sequence-length and amino-acid-composition baseline, while an InterPLM SAE at layer 18 failed preregistered stability criteria. We therefore examined the geometry of supervised toxin-related directions relative to the frozen InterPLM layer-18 decoder. Using orthogonal matching pursuit across biological, canonical label-null, C-matched null, and isotropic reference directions, 32 decoder atoms reconstructed a median 25.68% of the biological direction, compared with 24.33% for C-matched null directions and 24.12% for isotropic directions. The biological decoder-weight distribution had an effective support of approximately 6,374 of 10,231 usable decoder atoms. All biological directions required more than 64 atoms to reach 50% reconstruction, more than 128 to reach 80%, and more than 256 to reach 90%. The dominant result is therefore strong linear accessibility without compact SAE-basis decomposition. The earlier SAE stability failure cannot be attributed to absence of linearly accessible toxin-related information; instead, the distributed decoder geometry provides a plausible geometric explanation for why a compact recurrent feature set was not recovered. Modest preferential decoder alignment relative to controls is visible in point estimates, but biological bootstrap intervals primarily characterize solver-seed variation conditional on a fixed dataset and should not be interpreted as biological sampling uncertainty. These findings identify a contrasting regime to recent positive protein-SAE results and motivate further evaluation of when biological functions are sparsely localized versus distributed in protein foundation-model representations. The study is descriptive and discovery-only and does not establish cross-family generalization, causal biological mechanisms, or operational hazardous-sequence screening performance.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-09
DOI
https://doi.org/10.5281/zenodo.22681898
Primary Topic
Machine Learning in Bioinformatics
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Distributed, Not Sparse: Decoder-Basis Geometry of Toxin-Related Directions in a Protein Language Model

Allan Ochola
Zenodo (CERN European Organization for Nuclear Research)
Machine Learning in Bioinformatics
preprint

Distributed, Not Sparse: Decoder-Basis Geometry of Toxin-Related Directions in a Protein Language Model

Allan Ochola
preprint en

Abstract

Sparse autoencoders (SAEs) are increasingly used to identify interpretable features in protein language models, but a failure to recover stable sparse features can have at least two explanations: the relevant biological information may be absent from the underlying representation, or it may be present but poorly aligned with the learned sparse basis. This study distinguishes these possibilities for toxin-related information in ESM-2. Prior experiments found strong linear accessibility of toxin-related information in raw ESM-2 representations, with median held-out AUROC reaching 0.965 at layer 24 compared with 0.587 for a sequence-length and amino-acid-composition baseline, while an InterPLM SAE at layer 18 failed preregistered stability criteria. We therefore examined the geometry of supervised toxin-related directions relative to the frozen InterPLM layer-18 decoder. Using orthogonal matching pursuit across biological, canonical label-null, C-matched null, and isotropic reference directions, 32 decoder atoms reconstructed a median 25.68% of the biological direction, compared with 24.33% for C-matched null directions and 24.12% for isotropic directions. The biological decoder-weight distribution had an effective support of approximately 6,374 of 10,231 usable decoder atoms. All biological directions required more than 64 atoms to reach 50% reconstruction, more than 128 to reach 80%, and more than 256 to reach 90%. The dominant result is therefore strong linear accessibility without compact SAE-basis decomposition. The earlier SAE stability failure cannot be attributed to absence of linearly accessible toxin-related information; instead, the distributed decoder geometry provides a plausible geometric explanation for why a compact recurrent feature set was not recovered. Modest preferential decoder alignment relative to controls is visible in point estimates, but biological bootstrap intervals primarily characterize solver-seed variation conditional on a fixed dataset and should not be interpreted as biological sampling uncertainty. These findings identify a contrasting regime to recent positive protein-SAE results and motivate further evaluation of when biological functions are sparsely localized versus distributed in protein foundation-model representations. The study is descriptive and discovery-only and does not establish cross-family generalization, causal biological mechanisms, or operational hazardous-sequence screening performance.

Zenodo (CERN European Organization for Nuclear Research)
Machine Learning in Bioinformatics
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.