Selecting GPTQ Damping from the Calibration Set by Nested-Ridge Leave-One-Out
GPTQ-style quantization damps the calibration Hessian by a fixed fraction of its mean diagonal (one percent in the GPTQ authors' code, five percent in a maintained implementation), a choice that ignores how many calibration rows there are. Small tabular, time-series and text-embedding foundation models can be calibrated with about as many rows as their layers are wide, and there no fixed fraction fits: on the layers we measure, the best fixed multiplier moves from 0.0064 where rows and width are of the same order to 0.77 where calibration is scarcer. Building on the known damping–ridge identity, we show that the Cholesky factor GPTQ already builds holds nested ridge regressions, one per input coordinate, that the damping regularizes. Under an idealized rounding model, this reading decomposes the layer's population error exactly and bounds, from the population activation spectrum alone, what GPTQ can gain over rounding to nearest. Selecting each layer's penalty by the leave-one-out error summed over the regressions (LOOCV) needs no extra quantization pass, because that error comes out of the same factorization. On 175 real-layer cells from four foundation models, LOOCV, chosen among six cross-validation arms (LOOCV and five generalized cross-validation variants), has a median excess cost (extra error over the best evaluated error, as a fraction of the best error's gain over rounding to nearest) of 0.0048 against 0.0843 for the one-percent default and 0.0401 for the five-percent rule, lower than every fixed, searched or hindsight rule we compare; a confirmation on re-split calibration pools and a model absent from the selection repeats its advantage over the one-percent default where calibration rows and layer width are of the same order. Quantizing the Qwen3.5-0.8B language model whole to 4 bits, LOOCV gives the lowest mean WikiText-2 perplexity of the rules compared in each of three calibration-size and format settings. In the per-channel format the 95% intervals of its advantage over the five-percent rule, not corrected for multiplicity, are above zero with 512 calibration tokens and include zero with 2048 ([−0.0003, +0.0577]). Where calibration rows and layer width are of the same order, damping can therefore be estimated from each layer's calibration set rather than fixed.
Authors
- Ya-Fen Yeh
- Guan-Yuan Chen (ORCID: https://orcid.org/0000-0003-3298-0624)
Institutions
- National Tsing Hua University (TW)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-28
- DOI
- https://doi.org/10.5281/zenodo.23004086
- Primary Topic
- Scientific Computing and Data Management
- Type
- preprint