When Explanations Write Back: Mechanistic Provenance for Language Models
Mechanistic explanations of language models can become causally entangled with the systems they are intended to describe when explanation content is written back into writable model state before validation. This paper formalizes that problem as mechanistic provenance, proves that post-writeback fidelity need not identify fidelity to the pre-writeback mechanism, and introduces the Baseline Write Barrier: a version-preserving evaluation protocol that retains an untouched pre-explanation branch. Controlled experiments show that counterfactual report content can shift subsequently measured causal feature dependence. A preregistered SST-2-derived replication with Phi-3.5-mini-instruct yields a mapping-balanced report-meaning contrast of 0.03317 (95% CI [0.02493, 0.04202]; 87/96 families; exact sign-test p = 3.64×10^-17). The stronger preregistered cross-architecture gate was not met, and separate experiments did not support privileged introspection. This record also contains the complete frozen public reproducibility package. It includes the exact PR-002 executed script, all 96 Phi family-level validation records, original Hugging Face Jobs provenance and hashes, preregistered protocols, post hoc robustness analyses clearly separated from confirmatory results, integrity-verification scripts, and synchronized manuscript artifacts. The archive is self-contained for inspection of the PR-002 confirmatory record.
Authors
- Aidan Edward Lawson (ORCID: https://orcid.org/0009-0007-1839-4801)
Institutions
- Institute of Medical Ethics (GB)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-12
- DOI
- https://doi.org/10.5281/zenodo.22728293
- Primary Topic
- Scientific Computing and Data Management
- Type
- preprint