Accuracy-Matched Is Not Cost-Matched: Pricing and Repairing Token Inflation in INT4 Reasoning Models

INT4 weight quantization is accepted for a reasoning model when it preserves accuracy, but a reasoning model chooses how many tokens it generates, so an INT4 model that matches BF16 accuracy can still cost more per solved problem. Here we price a compressed reasoner by its generated tokens per solved problem as a paired ratio to BF16, test that ratio for equivalence in the form one would deploy, and time that form serving the whole test set end to end. On GSM8K, GPTQ INT4 costs 5.0–7.4% more per solved problem than BF16 in three of four open reasoning models while pass@1 stays within 0.4 points; a low-rank distillation adapter on the frozen INT4 weights brings all three within ±5% of BF16 (deployed Qwen3-4B: 1.016, 90% CI [1.001, 1.030]), whereas on the Qwen3-4B full test set hard-label training on the same data does not. On Qwen3-4B, a stronger post-training quantizer reaches the same token cost without an adapter. Tokens saved are not seconds saved: served unmerged in vLLM, the Qwen3-4B adapter adds 10.2–172.3% per token in a separate serving benchmark, and on an RTX 4090 it takes 1.212 times its INT4 base’s wall-clock time at batch 1 and 0.981 times (95% interval [0.963, 1.000]) when the full test set is served with up to 128 concurrent sequences, whereas merging the adapter and requantizing to INT4 takes 0.861 times with pass@1 within 0.01. Compressed reasoning models should therefore be evaluated by cost per solved problem in the deployed form, and a repair priced by that form’s end-to-end time under the load it will serve, not by per-token benchmarks alone.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-05
DOI
https://doi.org/10.5281/zenodo.23153532
Primary Topic
Advanced Neural Network Applications
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Accuracy-Matched Is Not Cost-Matched: Pricing and Repairing Token Inflation in INT4 Reasoning Models

Ya-Fen Yeh, Guan-Yuan Chen
Zenodo (CERN European Organization for Nuclear Research)
Advanced Neural Network Applications
preprint

Accuracy-Matched Is Not Cost-Matched: Pricing and Repairing Token Inflation in INT4 Reasoning Models

Ya-Fen Yeh, Guan-Yuan Chen
preprint en

Abstract

INT4 weight quantization is accepted for a reasoning model when it preserves accuracy, but a reasoning model chooses how many tokens it generates, so an INT4 model that matches BF16 accuracy can still cost more per solved problem. Here we price a compressed reasoner by its generated tokens per solved problem as a paired ratio to BF16, test that ratio for equivalence in the form one would deploy, and time that form serving the whole test set end to end. On GSM8K, GPTQ INT4 costs 5.0–7.4% more per solved problem than BF16 in three of four open reasoning models while pass@1 stays within 0.4 points; a low-rank distillation adapter on the frozen INT4 weights brings all three within ±5% of BF16 (deployed Qwen3-4B: 1.016, 90% CI [1.001, 1.030]), whereas on the Qwen3-4B full test set hard-label training on the same data does not. On Qwen3-4B, a stronger post-training quantizer reaches the same token cost without an adapter. Tokens saved are not seconds saved: served unmerged in vLLM, the Qwen3-4B adapter adds 10.2–172.3% per token in a separate serving benchmark, and on an RTX 4090 it takes 1.212 times its INT4 base’s wall-clock time at batch 1 and 0.981 times (95% interval [0.963, 1.000]) when the full test set is served with up to 128 concurrent sequences, whereas merging the adapter and requantizing to INT4 takes 0.861 times with pass@1 within 0.01. Compressed reasoning models should therefore be evaluated by cost per solved problem in the deployed form, and a repair priced by that form’s end-to-end time under the load it will serve, not by per-token benchmarks alone.

Zenodo (CERN European Organization for Nuclear Research)
National Tsing Hua University (TW)
Advanced Neural Network Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.