AI Mechanistic Interpretability and Self-Aware Networks: From Neural Decoding to Causal Self-Regulation

An internal representation can be decodable, causally effective and numerically reconstructable without implementing the operation a user intended. I examine these distinctions through the receiver-relative framework of Self-Aware Networks, connecting a source-recovered neural-decoding and feedback question to explicit transformer interventions. The research combines intervention prediction in controlled learned networks, passive and active feedback controls, GPT-2 activation and cache experiments, and native hybrid-state interventions in a pinned quantized Qwen3.5-0.8B decoder. Complete corrected-state transfer supports all eight tested bidirectional story groups in an initial Qwen assay, while attention-only transfer fails one; a complementary intervention gives identical attention state with different behavior. In a separate learned-correction assay, linear and nonlinear attention-state predictors each obtain five of 32 requested reversed-role answers but no story with both roles correctly reversed. Both preserve the tested color attribute. Lower training loss and partial-state error therefore fail to certify the intended relational change. Formal arguments distinguish observational from causal identification, partial-state from behavioral sufficiency, and output-margin guarantees from reconstruction loss. Earlier failures, nulls and inadequate baselines are retained. The contribution is a bounded experimental account of receiver-dependent intervention and its certification problem, not invention of mechanistic interpretability, evidence of phenomenal consciousness, or demonstrated autonomous self-regulation in frontier models. Preprint v1.0; integrated Draft 10, September 5, 2026. The checksum-linked shared companion, archived with DOI 10.5281/zenodo.22360017, includes the applications, protocols, data, figures, 31 accepted Lean statements under recorded assumptions, and a later full fixed learned-Qwen pipeline reproduction (49 serial steps; 465 original answers and 465 separate-code reconstructions). The same experiments support both AI papers; these are not independent replications. The negative learned-correction result is retained. Authoring-agent checks and same-workflow reconstruction are not independent scientific review. Base-model and runtime binaries and private source originals are not included. License scope: CC BY 4.0 for original manuscript prose, figures, documentation and data. Original code is public for review without a separate reuse license; third-party materials retain their own terms. See the companion LICENSE.md and notices.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-05
DOI
https://doi.org/10.5281/zenodo.22360027
Primary Topic
Explainable Artificial Intelligence (XAI)
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

AI Mechanistic Interpretability and Self-Aware Networks: From Neural Decoding to Causal Self-Regulation

Micah Blumberg
Zenodo (CERN European Organization for Nuclear Research)
Explainable Artificial Intelligence (XAI)
preprint

AI Mechanistic Interpretability and Self-Aware Networks: From Neural Decoding to Causal Self-Regulation

Micah Blumberg
preprint en

Abstract

An internal representation can be decodable, causally effective and numerically reconstructable without implementing the operation a user intended. I examine these distinctions through the receiver-relative framework of Self-Aware Networks, connecting a source-recovered neural-decoding and feedback question to explicit transformer interventions. The research combines intervention prediction in controlled learned networks, passive and active feedback controls, GPT-2 activation and cache experiments, and native hybrid-state interventions in a pinned quantized Qwen3.5-0.8B decoder. Complete corrected-state transfer supports all eight tested bidirectional story groups in an initial Qwen assay, while attention-only transfer fails one; a complementary intervention gives identical attention state with different behavior. In a separate learned-correction assay, linear and nonlinear attention-state predictors each obtain five of 32 requested reversed-role answers but no story with both roles correctly reversed. Both preserve the tested color attribute. Lower training loss and partial-state error therefore fail to certify the intended relational change. Formal arguments distinguish observational from causal identification, partial-state from behavioral sufficiency, and output-margin guarantees from reconstruction loss. Earlier failures, nulls and inadequate baselines are retained. The contribution is a bounded experimental account of receiver-dependent intervention and its certification problem, not invention of mechanistic interpretability, evidence of phenomenal consciousness, or demonstrated autonomous self-regulation in frontier models. Preprint v1.0; integrated Draft 10, September 5, 2026. The checksum-linked shared companion, archived with DOI 10.5281/zenodo.22360017, includes the applications, protocols, data, figures, 31 accepted Lean statements under recorded assumptions, and a later full fixed learned-Qwen pipeline reproduction (49 serial steps; 465 original answers and 465 separate-code reconstructions). The same experiments support both AI papers; these are not independent replications. The negative learned-correction result is retained. Authoring-agent checks and same-workflow reconstruction are not independent scientific review. Base-model and runtime binaries and private source originals are not included. License scope: CC BY 4.0 for original manuscript prose, figures, documentation and data. Original code is public for review without a separate reuse license; third-party materials retain their own terms. See the companion LICENSE.md and notices.

Zenodo (CERN European Organization for Nuclear Research)
Kitware (United States) (US)
Peace, Justice and strong institutions
Explainable Artificial Intelligence (XAI)
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.