A Question Worth Answering: What Would a Neural Network Preserve?
The report asks a simple but focused question: if a numerical mechanism attached to a trained neural network had to select exactly one component to preserve, what would it choose, and would that choice predict real functional consequences better than simple baselines? We ran an exploratory series of experiments on small ReLU and tanh MLPs (breast-cancer dataset, then MNIST and Fashion-MNIST) and a small CNN. We tested internal importance criteria (reconstruction error, activity RMS/SD, gradient, ridge/peer prediction, perturbation stability), measured how well they ranked units or channels by ablation damage, and examined group-removal effects on accuracy and cross-entropy. Protocols, candidate definitions, and training scope changed across stages; all results come from the supplied logs and were not independently rerun. Claims that can be made: Internal importance is conditional on criterion, layer, architecture, candidate type (unit vs channel), dataset, and evaluation endpoint. Within-layer ranking associations, multi-unit removal damage, and actual preservation of capability are distinct and should not be conflated. Coupled internal rewards can be recovered, but this does not by itself prove an inspector-specific advantage. No universal preservation rule, intrinsic “intention,” or general superiority over simple baselines is established. Key data: Fifteen-seed equal-width MLP follow-up: residual reconstruction–damage partial rank correlations of 0.5395, 0.2711, and 0.2704 after adjusting for activity RMS and SD ranks (marginal seed-bootstrap intervals exclude zero). Eight-seed group removal (13 units ≈ 20 %): largest mean accuracy drops consistently appeared in the first layer on both MNIST and Fashion-MNIST for the tested criteria; advantages over random were common but not universal. Five-seed CNN channel check: single-channel correlations and removal layer order mixed across MNIST vs Fashion-MNIST and did not cleanly transfer the MLP patterns. Early stability and ridge selection phases produced 0 exact oracle matches under the reported conditions. Parts of the brainstorming, code scaffolding, analysis, and drafting used large language models (ChatGPT, Gemini, Grok, Claude). The author selected, edited, and remains responsible for all methods, results, and interpretations. Links to the originating conversations and the experiment Colab notebook are supplied in the report.
Authors
- Sohan Poudel (ORCID: https://orcid.org/0009-0004-5762-3898)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-13
- DOI
- https://doi.org/10.5281/zenodo.22735506
- Primary Topic
- Explainable Artificial Intelligence (XAI)
- Type
- preprint