Selecting GPTQ Damping from the Calibration Set by Nested-Ridge Leave-One-Out

GPTQ-style quantization damps the calibration Hessian by a fixed fraction of its mean diagonal (one percent in the GPTQ authors' code, five percent in a maintained implementation), a choice that ignores how many calibration rows there are. Small tabular, time-series and text-embedding foundation models can be calibrated with about as many rows as their layers are wide or fewer, and there no fixed fraction fits: on the layers we measure, the best fixed multiplier moves from 0.0064 where rows and width are of the same order to 0.77 where calibration is scarcer. Building on the known damping–ridge identity, we show that the Cholesky factor GPTQ already builds holds the d nested ridge regressions that the damping regularizes. Under an idealized rounding model, this reading decomposes the layer's population error exactly and bounds, from the population activation spectrum alone, what GPTQ can gain over rounding to nearest; it also makes the penalty estimable, because the leave-one-out error summed over the regressions (LOOCV) comes out of the same factorization with no extra quantization pass. On 175 real-layer cells with at most 512 calibration rows from four foundation models, LOOCV, chosen among six cross-validation criteria, has a median excess cost (extra error over the best error among the selection sweep's arms, as a fraction of its gain over rounding to nearest) of 0.0048 against 0.0843 for the one-percent default and 0.0401 for the five-percent rule, lower than every fixed, searched or hindsight rule we compare; a confirmation with the rule fixed in advance, on re-split calibration pools and a model absent from the selection, repeats its advantage over the one-percent default where calibration rows and layer width are of the same order. In whole-model 4-bit quantization of the Qwen3.5-0.8B language model, LOOCV also gives lower WikiText-2 perplexity than the one-percent default in each of the three calibration-size and format settings we test, and lower than the five-percent rule with 512 calibration tokens. Where calibration rows and layer width are of the same order, damping can therefore be estimated from each layer's calibration set rather than fixed.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-28
DOI
https://doi.org/10.5281/zenodo.23015545
Primary Topic
Scientific Computing and Data Management
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Selecting GPTQ Damping from the Calibration Set by Nested-Ridge Leave-One-Out

Ya-Fen Yeh, Guan-Yuan Chen
Zenodo (CERN European Organization for Nuclear Research)
Scientific Computing and Data Management
preprint

Selecting GPTQ Damping from the Calibration Set by Nested-Ridge Leave-One-Out

Ya-Fen Yeh, Guan-Yuan Chen
preprint en

Abstract

GPTQ-style quantization damps the calibration Hessian by a fixed fraction of its mean diagonal (one percent in the GPTQ authors' code, five percent in a maintained implementation), a choice that ignores how many calibration rows there are. Small tabular, time-series and text-embedding foundation models can be calibrated with about as many rows as their layers are wide or fewer, and there no fixed fraction fits: on the layers we measure, the best fixed multiplier moves from 0.0064 where rows and width are of the same order to 0.77 where calibration is scarcer. Building on the known damping–ridge identity, we show that the Cholesky factor GPTQ already builds holds the d nested ridge regressions that the damping regularizes. Under an idealized rounding model, this reading decomposes the layer's population error exactly and bounds, from the population activation spectrum alone, what GPTQ can gain over rounding to nearest; it also makes the penalty estimable, because the leave-one-out error summed over the regressions (LOOCV) comes out of the same factorization with no extra quantization pass. On 175 real-layer cells with at most 512 calibration rows from four foundation models, LOOCV, chosen among six cross-validation criteria, has a median excess cost (extra error over the best error among the selection sweep's arms, as a fraction of its gain over rounding to nearest) of 0.0048 against 0.0843 for the one-percent default and 0.0401 for the five-percent rule, lower than every fixed, searched or hindsight rule we compare; a confirmation with the rule fixed in advance, on re-split calibration pools and a model absent from the selection, repeats its advantage over the one-percent default where calibration rows and layer width are of the same order. In whole-model 4-bit quantization of the Qwen3.5-0.8B language model, LOOCV also gives lower WikiText-2 perplexity than the one-percent default in each of the three calibration-size and format settings we test, and lower than the five-percent rule with 512 calibration tokens. Where calibration rows and layer width are of the same order, damping can therefore be estimated from each layer's calibration set rather than fixed.

Zenodo (CERN European Organization for Nuclear Research)
National Tsing Hua University (TW)
Quality Education
Scientific Computing and Data Management
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.