Auditing the Transferability of a Published Quantization Comparison: The Verdict Is a Function of the Slice

Practitioners choose between released 4-bit conversions of a model from published comparison sentences such as "AWQ consistently outperforms ... GPTQ ... (7B-70B)", although the files they serve, the serving stack and the way quality is measured all differ from what the sentence was written about. Later comparisons disagree about its direction without reporting intervals, so whether such a verdict transfers to deployment artifacts is unknown. We test the transfer on released artifact pairs by crossing the measurement regime (perplexity against downstream accuracy) with the serving engine on one 12 GB accelerator, and report each verdict with its interval, margin and realized precision. On an in-range 7B released pair from a family outside the audited table, WikiText-2 perplexity places AWQ behind by 0.993% while LAMBADA accuracy, chosen after reading eight tasks, places it ahead by 3.136 points, both beyond their margins on both engines; the opposition persists when one dataset is scored both ways. In a five-cell verdict matrix the measurement regime changed verdict states and the serving engine none; where the engines disagreed elsewhere (one audited-table checkpoint; BoolQ at 1.5B and 7B; LAMBADA on a preliminary second family), the intervals split between inconclusive and zero-excluded, and on the primary family's two pairs most tasks fixed before measurement, the primary one included, did not resolve their margin. The direction of a published quantization comparison can thus depend on the slice it is read on, so each verdict needs its slice and realized precision stated beside it.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-25
DOI
https://doi.org/10.5281/zenodo.22961364
Primary Topic
Scientific Computing and Data Management
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Auditing the Transferability of a Published Quantization Comparison: The Verdict Is a Function of the Slice

Ya-Fen Yeh, Guan-Yuan Chen
Zenodo (CERN European Organization for Nuclear Research)
Scientific Computing and Data Management
preprint

Auditing the Transferability of a Published Quantization Comparison: The Verdict Is a Function of the Slice

Ya-Fen Yeh, Guan-Yuan Chen
preprint en

Abstract

Practitioners choose between released 4-bit conversions of a model from published comparison sentences such as "AWQ consistently outperforms ... GPTQ ... (7B-70B)", although the files they serve, the serving stack and the way quality is measured all differ from what the sentence was written about. Later comparisons disagree about its direction without reporting intervals, so whether such a verdict transfers to deployment artifacts is unknown. We test the transfer on released artifact pairs by crossing the measurement regime (perplexity against downstream accuracy) with the serving engine on one 12 GB accelerator, and report each verdict with its interval, margin and realized precision. On an in-range 7B released pair from a family outside the audited table, WikiText-2 perplexity places AWQ behind by 0.993% while LAMBADA accuracy, chosen after reading eight tasks, places it ahead by 3.136 points, both beyond their margins on both engines; the opposition persists when one dataset is scored both ways. In a five-cell verdict matrix the measurement regime changed verdict states and the serving engine none; where the engines disagreed elsewhere (one audited-table checkpoint; BoolQ at 1.5B and 7B; LAMBADA on a preliminary second family), the intervals split between inconclusive and zero-excluded, and on the primary family's two pairs most tasks fixed before measurement, the primary one included, did not resolve their margin. The direction of a published quantization comparison can thus depend on the slice it is read on, so each verdict needs its slice and realized precision stated beside it.

Zenodo (CERN European Organization for Nuclear Research)
National Tsing Hua University (TW), North Carolina Exploring Cultural Heritage Online (US)
Peace, Justice and strong institutions
Scientific Computing and Data Management
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Auditing the Transferability of a Published Quantization Comparison: The Verdict Is a Function of the Slice — Ya-Fen Yeh, Guan-Yuan Chen · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS