Arithmetic Performance After Transformer Layer Ablation: A Case Study of Qwen2.5-1.5B-Instruct

Removing transformer blocks can preserve substantial task performance, but aggregate accuracy can conceal changes in numerical answers, option selection, and response formatting. I investigate these outcomes in Qwen2.5-1.5B-Instruct using arithmetic problems presented with and without answer choices. An exploratory layer sweep identifies layer 4 removal as a promising condition, which I then examine with a larger generation budget and a held-out test set. Allowing longer responses narrows the accuracy gap but does not eliminate it. On the held-out test set of 120 problems presented in three formats, removing layer 4 improves final-answer accuracy from 87.2% to 91.1%, a gain of 3.89 percentage points (paired problem-cluster bootstrap 95% interval: 0.56 to 7.50). The gain is concentrated in multiple-choice prompts, while no-option accuracy declines and responses become substantially shorter. These findings show how separating numerical conclusions, option selection, and formatting helps explain an apparent improvement after layer ablation. This manuscript is a preprint and has not undergone peer review.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-05
DOI
https://doi.org/10.5281/zenodo.23154624
Primary Topic
Topic Modeling
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Arithmetic Performance After Transformer Layer Ablation: A Case Study of Qwen2.5-1.5B-Instruct

Sahil Basera
Zenodo (CERN European Organization for Nuclear Research)
Topic Modeling
preprint

Arithmetic Performance After Transformer Layer Ablation: A Case Study of Qwen2.5-1.5B-Instruct

Sahil Basera
preprint en

Abstract

Removing transformer blocks can preserve substantial task performance, but aggregate accuracy can conceal changes in numerical answers, option selection, and response formatting. I investigate these outcomes in Qwen2.5-1.5B-Instruct using arithmetic problems presented with and without answer choices. An exploratory layer sweep identifies layer 4 removal as a promising condition, which I then examine with a larger generation budget and a held-out test set. Allowing longer responses narrows the accuracy gap but does not eliminate it. On the held-out test set of 120 problems presented in three formats, removing layer 4 improves final-answer accuracy from 87.2% to 91.1%, a gain of 3.89 percentage points (paired problem-cluster bootstrap 95% interval: 0.56 to 7.50). The gain is concentrated in multiple-choice prompts, while no-option accuracy declines and responses become substantially shorter. These findings show how separating numerical conclusions, option selection, and formatting helps explain an apparent improvement after layer ablation. This manuscript is a preprint and has not undergone peer review.

Zenodo (CERN European Organization for Nuclear Research)
Topic Modeling
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Arithmetic Performance After Transformer Layer Ablation: A Case Study of Qwen2.5-1.5B-Instruct — Sahil Basera · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS