Mechanistic Interpretability and Explainable AI (XAI) in Foundation Models: Reverse-Engineering Circuits, Superposition, and Safety Bounds
Abstract—Foundation models, particularly autoregressive Transformer Large Language Models (LLMs) and Vision-Language Models (VLMs), have exhibited unprecedented reasoning and generative capabilities across computer science and engineering. However, their internal representations operate as opaque, high-dimensional 'black boxes.' Conventional Explainable AI (XAI) techniques—predominantly post-hoc attribution methods such as Integrated Gradients, SHAP, and attention heatmap visualizations—fail to provide faithful causal explanations: they reflect output correlations rather than the internal mechanistic computations executing within the network, masking hazardous failure modes like latent hallucinations, sycophancy, and deceptive alignment. Mechanistic Interpretability has emerged as a rigorous bottom-up paradigm that seeks to reverse-engineer deep neural networks into human-understandable algorithmic circuits, decomposing polysemantic neuron activations into monosemantic features. This paper provides an exhaustive, mathematically formal technical investigation of mechanistic interpretability and XAI for modern foundation models. We formalize the Linear Representation Hypothesis and the mathematics of neural superposition, analyze transformer residual streams as communicative communication buses, and model the discovery of specialized functional subgraphs—including induction heads, factual recall circuits, and indirect object identification (IOI) subnetworks. Furthermore, we evaluate Sparse Autoencoders (SAEs) as an unsupervised mechanism to disentangle polysemantic latent spaces into overcomplete, monosemantic feature dictionaries. Finally, we implement an empirical evaluation testbed across open foundation models (Llama-3-8B, Gemma-2-9B, and Mistral-7B), quantitatively benchmarking causal scrubbing, activation patching, and SAE feature steering across mission-critical tasks (deceptive alignment detection, factual hallucination suppression, and jailbreak defense). Our findings demonstrate that activation patching isolates causal circuits responsible for 92.4% of factual routing, while SAE-guided feature clamping suppresses unfaithful hallucinations by 41.7% with less than 2.1% degradation in baseline perplexity. We conclude by formulating open challenges in circuit universality, superposition scaling laws, and regulatory verification under international AI governance frameworks. Index Terms—Mechanistic Interpretability, Explainable AI (XAI), Foundation Models, Sparse Autoencoders (SAEs), Polysemanticity, Neural Superposition, Induction Heads, Activation Patching, Transformer Circuits.
Authors
- Megha Rathore
- Vinay Kumar Patidar
- Dr. Akhil Pandey
- Dr. Vishal Shrivastava
- Rashika
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-09
- DOI
- https://doi.org/10.5281/zenodo.23259022
- Primary Topic
- Explainable Artificial Intelligence (XAI)
- Type
- article
- Field-Weighted Citation Impact
- 0.00