Accuracy-Matched Is Not Cost-Matched: Pricing and Repairing Token Inflation in INT4 Reasoning Models
INT4 weight quantization is accepted for a reasoning model when it preserves accuracy, but a reasoning model chooses how many tokens it generates, so an INT4 model that matches BF16 accuracy can still cost more per solved problem. Here we price a compressed reasoner by its generated tokens per solved problem as a paired ratio to BF16, test that ratio for equivalence in the form one would deploy, and convert it to GPU time with measured per-token time. On GSM8K, GPTQ INT4 costs 5.0–7.4% more per solved problem than BF16 in three of four open reasoning models while pass@1 stays within 0.4 points; a low-rank distillation adapter on the frozen INT4 weights brings all three within ±5% of BF16 (deployed Qwen3-4B: 1.016, 90% CI [1.001, 1.030]), whereas on the Qwen3-4B full test set hard-label training on the same data does not. On Qwen3-4B, a stronger post-training quantizer reaches the same token cost without an adapter, and on the long-horizon MATH-500 subset the repair leaves the cost significantly above BF16. Served on an RTX 4090 in vLLM, however, the unmerged Qwen3-4B adapter adds 10.3–132.1% per token, above the 5.5% break-even set by its 5.2% token saving, at batch 1 and at saturation in both serving modes, so the repaired model costs more GPU time per solved problem than its INT4 base, and the stronger quantizer costs the least. Compressed reasoning models should therefore be evaluated by cost per solved problem in the deployed form, and a repair priced by the per-token time it adds as well as by the tokens it saves.
Authors
- Ya-Fen Yeh
- Guan-Yuan Chen (ORCID: https://orcid.org/0000-0003-3298-0624)
Institutions
- National Tsing Hua University (TW)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-28
- DOI
- https://doi.org/10.5281/zenodo.23014673
- Primary Topic
- Machine Learning and Data Classification
- Type
- preprint