Arithmetic Performance After Transformer Layer Ablation: A Case Study of Qwen2.5-1.5B-Instruct
Removing transformer blocks can preserve substantial task performance, but aggregate accuracy can conceal changes in numerical answers, option selection, and response formatting. I investigate these outcomes in Qwen2.5-1.5B-Instruct using arithmetic problems presented with and without answer choices. An exploratory layer sweep identifies layer 4 removal as a promising condition, which I then examine with a larger generation budget and a held-out test set. Allowing longer responses narrows the accuracy gap but does not eliminate it. On the held-out test set of 120 problems presented in three formats, removing layer 4 improves final-answer accuracy from 87.2% to 91.1%, a gain of 3.89 percentage points (paired problem-cluster bootstrap 95% interval: 0.56 to 7.50). The gain is concentrated in multiple-choice prompts, while no-option accuracy declines and responses become substantially shorter. These findings show how separating numerical conclusions, option selection, and formatting helps explain an apparent improvement after layer ablation. This manuscript is a preprint and has not undergone peer review.
Authors
- Sahil Basera
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-05
- DOI
- https://doi.org/10.5281/zenodo.23154625
- Primary Topic
- Topic Modeling
- Type
- preprint