Accuracy-Matched Is Not Cost-Matched: Pricing and Repairing Token Inflation in INT4 Reasoning Models
INT4 weight quantization is accepted for a reasoning model when it preserves accuracy, but a reasoning model chooses how many tokens it generates, so an INT4 model that matches BF16 accuracy can still cost more per solved problem. Here we price a compressed reasoner by its generated tokens per solved problem as a paired ratio to BF16, test that ratio for equivalence in the form one would deploy, and time that form serving the whole test set end to end. On GSM8K, GPTQ INT4 costs 5.0–7.4% more per solved problem than BF16 in three of four open reasoning models while pass@1 stays within 0.4 points; a low-rank distillation adapter on the frozen INT4 weights brings all three within ±5% of BF16 (deployed Qwen3-4B: 1.016, 90% CI [1.001, 1.030]), whereas on the Qwen3-4B full test set hard-label training on the same data does not. On Qwen3-4B, a stronger post-training quantizer reaches the same token cost without an adapter. Tokens saved are not seconds saved: served unmerged in vLLM, the Qwen3-4B adapter adds 10.2–172.3% per token in a separate serving benchmark, and on an RTX 4090 it takes 1.212 times its INT4 base’s wall-clock time at batch 1 and 0.981 times (95% interval [0.963, 1.000]) when the full test set is served with up to 128 concurrent sequences, whereas merging the adapter and requantizing to INT4 takes 0.861 times with pass@1 within 0.01. Compressed reasoning models should therefore be evaluated by cost per solved problem in the deployed form, and a repair priced by that form’s end-to-end time under the load it will serve, not by per-token benchmarks alone.
Authors
- Ya-Fen Yeh
- Guan-Yuan Chen (ORCID: https://orcid.org/0000-0003-3298-0624)
Institutions
- National Tsing Hua University (TW)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-05
- DOI
- https://doi.org/10.5281/zenodo.23153532
- Primary Topic
- Advanced Neural Network Applications
- Type
- preprint