A Question Worth Answering: What Would a Neural Network Preserve?

The report asks a simple but focused question: if a numerical mechanism attached to a trained neural network had to select exactly one component to preserve, what would it choose, and would that choice predict real functional consequences better than simple baselines? We ran an exploratory series of experiments on small ReLU and tanh MLPs (breast-cancer dataset, then MNIST and Fashion-MNIST) and a small CNN. We tested internal importance criteria (reconstruction error, activity RMS/SD, gradient, ridge/peer prediction, perturbation stability), measured how well they ranked units or channels by ablation damage, and examined group-removal effects on accuracy and cross-entropy. Protocols, candidate definitions, and training scope changed across stages; all results come from the supplied logs and were not independently rerun. Claims that can be made: Internal importance is conditional on criterion, layer, architecture, candidate type (unit vs channel), dataset, and evaluation endpoint. Within-layer ranking associations, multi-unit removal damage, and actual preservation of capability are distinct and should not be conflated. Coupled internal rewards can be recovered, but this does not by itself prove an inspector-specific advantage. No universal preservation rule, intrinsic “intention,” or general superiority over simple baselines is established. Key data: Fifteen-seed equal-width MLP follow-up: residual reconstruction–damage partial rank correlations of 0.5395, 0.2711, and 0.2704 after adjusting for activity RMS and SD ranks (marginal seed-bootstrap intervals exclude zero). Eight-seed group removal (13 units ≈ 20 %): largest mean accuracy drops consistently appeared in the first layer on both MNIST and Fashion-MNIST for the tested criteria; advantages over random were common but not universal. Five-seed CNN channel check: single-channel correlations and removal layer order mixed across MNIST vs Fashion-MNIST and did not cleanly transfer the MLP patterns. Early stability and ridge selection phases produced 0 exact oracle matches under the reported conditions. Parts of the brainstorming, code scaffolding, analysis, and drafting used large language models (ChatGPT, Gemini, Grok, Claude). The author selected, edited, and remains responsible for all methods, results, and interpretations. Links to the originating conversations and the experiment Colab notebook are supplied in the report.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-13
DOI
https://doi.org/10.5281/zenodo.22735506
Primary Topic
Explainable Artificial Intelligence (XAI)
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

A Question Worth Answering: What Would a Neural Network Preserve?

Sohan Poudel
Zenodo (CERN European Organization for Nuclear Research)
Explainable Artificial Intelligence (XAI)
preprint

A Question Worth Answering: What Would a Neural Network Preserve?

Sohan Poudel
preprint en

Abstract

The report asks a simple but focused question: if a numerical mechanism attached to a trained neural network had to select exactly one component to preserve, what would it choose, and would that choice predict real functional consequences better than simple baselines? We ran an exploratory series of experiments on small ReLU and tanh MLPs (breast-cancer dataset, then MNIST and Fashion-MNIST) and a small CNN. We tested internal importance criteria (reconstruction error, activity RMS/SD, gradient, ridge/peer prediction, perturbation stability), measured how well they ranked units or channels by ablation damage, and examined group-removal effects on accuracy and cross-entropy. Protocols, candidate definitions, and training scope changed across stages; all results come from the supplied logs and were not independently rerun. Claims that can be made: Internal importance is conditional on criterion, layer, architecture, candidate type (unit vs channel), dataset, and evaluation endpoint. Within-layer ranking associations, multi-unit removal damage, and actual preservation of capability are distinct and should not be conflated. Coupled internal rewards can be recovered, but this does not by itself prove an inspector-specific advantage. No universal preservation rule, intrinsic “intention,” or general superiority over simple baselines is established. Key data: Fifteen-seed equal-width MLP follow-up: residual reconstruction–damage partial rank correlations of 0.5395, 0.2711, and 0.2704 after adjusting for activity RMS and SD ranks (marginal seed-bootstrap intervals exclude zero). Eight-seed group removal (13 units ≈ 20 %): largest mean accuracy drops consistently appeared in the first layer on both MNIST and Fashion-MNIST for the tested criteria; advantages over random were common but not universal. Five-seed CNN channel check: single-channel correlations and removal layer order mixed across MNIST vs Fashion-MNIST and did not cleanly transfer the MLP patterns. Early stability and ridge selection phases produced 0 exact oracle matches under the reported conditions. Parts of the brainstorming, code scaffolding, analysis, and drafting used large language models (ChatGPT, Gemini, Grok, Claude). The author selected, edited, and remains responsible for all methods, results, and interpretations. Links to the originating conversations and the experiment Colab notebook are supplied in the report.

Zenodo (CERN European Organization for Nuclear Research)
Explainable Artificial Intelligence (XAI)
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.