When Explanations Write Back: Mechanistic Provenance for Language Models

Mechanistic explanations of language models can become causally entangled with the systems they are intended to describe when explanation content is written back into writable model state before validation. This paper formalizes that problem as mechanistic provenance, proves that post-writeback fidelity need not identify fidelity to the pre-writeback mechanism, and introduces the Baseline Write Barrier: a version-preserving evaluation protocol that retains an untouched pre-explanation branch. Controlled experiments show that counterfactual report content can shift subsequently measured causal feature dependence. A preregistered SST-2-derived replication with Phi-3.5-mini-instruct yields a mapping-balanced report-meaning contrast of 0.03317 (95% CI [0.02493, 0.04202]; 87/96 families; exact sign-test p = 3.64×10^-17). The stronger preregistered cross-architecture gate was not met, and separate experiments did not support privileged introspection. This record also contains the complete frozen public reproducibility package. It includes the exact PR-002 executed script, all 96 Phi family-level validation records, original Hugging Face Jobs provenance and hashes, preregistered protocols, post hoc robustness analyses clearly separated from confirmatory results, integrity-verification scripts, and synchronized manuscript artifacts. The archive is self-contained for inspection of the PR-002 confirmatory record.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-12
DOI
https://doi.org/10.5281/zenodo.22728293
Primary Topic
Scientific Computing and Data Management
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

When Explanations Write Back: Mechanistic Provenance for Language Models

Aidan Edward Lawson
Zenodo (CERN European Organization for Nuclear Research)
Scientific Computing and Data Management
preprint

When Explanations Write Back: Mechanistic Provenance for Language Models

Aidan Edward Lawson
preprint en

Abstract

Mechanistic explanations of language models can become causally entangled with the systems they are intended to describe when explanation content is written back into writable model state before validation. This paper formalizes that problem as mechanistic provenance, proves that post-writeback fidelity need not identify fidelity to the pre-writeback mechanism, and introduces the Baseline Write Barrier: a version-preserving evaluation protocol that retains an untouched pre-explanation branch. Controlled experiments show that counterfactual report content can shift subsequently measured causal feature dependence. A preregistered SST-2-derived replication with Phi-3.5-mini-instruct yields a mapping-balanced report-meaning contrast of 0.03317 (95% CI [0.02493, 0.04202]; 87/96 families; exact sign-test p = 3.64×10^-17). The stronger preregistered cross-architecture gate was not met, and separate experiments did not support privileged introspection. This record also contains the complete frozen public reproducibility package. It includes the exact PR-002 executed script, all 96 Phi family-level validation records, original Hugging Face Jobs provenance and hashes, preregistered protocols, post hoc robustness analyses clearly separated from confirmatory results, integrity-verification scripts, and synchronized manuscript artifacts. The archive is self-contained for inspection of the PR-002 confirmatory record.

Zenodo (CERN European Organization for Nuclear Research)
Institute of Medical Ethics (GB)
Scientific Computing and Data Management
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

When Explanations Write Back: Mechanistic Provenance for Language Models — Aidan Edward Lawson · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS