Accuracy-Matched Is Not Cost-Matched: Pricing and Repairing Token Inflation in INT4 Reasoning Models

INT4 weight quantization is accepted for a reasoning model when it preserves accuracy, but a reasoning model chooses how many tokens it generates, so an INT4 model that matches BF16 accuracy can still cost more per solved problem. Here we price a compressed reasoner by its generated tokens per solved problem as a paired ratio to BF16, test that ratio for equivalence in the form one would deploy, and convert it to GPU time with measured per-token time. On GSM8K, GPTQ INT4 costs 5.0–7.4% more per solved problem than BF16 in three of four open reasoning models while pass@1 stays within 0.4 points; a low-rank distillation adapter on the frozen INT4 weights brings all three within ±5% of BF16 (deployed Qwen3-4B: 1.016, 90% CI [1.001, 1.030]), whereas on the Qwen3-4B full test set hard-label training on the same data does not. On Qwen3-4B, a stronger post-training quantizer reaches the same token cost without an adapter, and on the long-horizon MATH-500 subset the repair leaves the cost significantly above BF16. Served on an RTX 4090 in vLLM, however, the unmerged Qwen3-4B adapter adds 10.3–132.1% per token, above the 5.5% break-even set by its 5.2% token saving, at batch 1 and at saturation in both serving modes, so the repaired model costs more GPU time per solved problem than its INT4 base, and the stronger quantizer costs the least. Compressed reasoning models should therefore be evaluated by cost per solved problem in the deployed form, and a repair priced by the per-token time it adds as well as by the tokens it saves.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-28
DOI
https://doi.org/10.5281/zenodo.22670367
Primary Topic
Constraint Satisfaction and Optimization
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Accuracy-Matched Is Not Cost-Matched: Pricing and Repairing Token Inflation in INT4 Reasoning Models

Ya-Fen Yeh, Guan-Yuan Chen
Zenodo (CERN European Organization for Nuclear Research)
Constraint Satisfaction and Optimization
preprint

Accuracy-Matched Is Not Cost-Matched: Pricing and Repairing Token Inflation in INT4 Reasoning Models

Ya-Fen Yeh, Guan-Yuan Chen
preprint en

Abstract

INT4 weight quantization is accepted for a reasoning model when it preserves accuracy, but a reasoning model chooses how many tokens it generates, so an INT4 model that matches BF16 accuracy can still cost more per solved problem. Here we price a compressed reasoner by its generated tokens per solved problem as a paired ratio to BF16, test that ratio for equivalence in the form one would deploy, and convert it to GPU time with measured per-token time. On GSM8K, GPTQ INT4 costs 5.0–7.4% more per solved problem than BF16 in three of four open reasoning models while pass@1 stays within 0.4 points; a low-rank distillation adapter on the frozen INT4 weights brings all three within ±5% of BF16 (deployed Qwen3-4B: 1.016, 90% CI [1.001, 1.030]), whereas on the Qwen3-4B full test set hard-label training on the same data does not. On Qwen3-4B, a stronger post-training quantizer reaches the same token cost without an adapter, and on the long-horizon MATH-500 subset the repair leaves the cost significantly above BF16. Served on an RTX 4090 in vLLM, however, the unmerged Qwen3-4B adapter adds 10.3–132.1% per token, above the 5.5% break-even set by its 5.2% token saving, at batch 1 and at saturation in both serving modes, so the repaired model costs more GPU time per solved problem than its INT4 base, and the stronger quantizer costs the least. Compressed reasoning models should therefore be evaluated by cost per solved problem in the deployed form, and a repair priced by the per-token time it adds as well as by the tokens it saves.

Zenodo (CERN European Organization for Nuclear Research)
National Tsing Hua University (TW), North Carolina Exploring Cultural Heritage Online (US)
Decent work and economic growth
Constraint Satisfaction and Optimization
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.