Same Accuracy, Different Verdicts: Quantization-Induced Verdict Churn in Small LLM Smart-Contract Auditors

Small open LLMs are increasingly run as smart-contract auditors in quantized GGUF form, and large models are reached through API aggregators that do not disclose how they are served; both practices are usually justified by aggregate accuracy. We measure verdict churn, the share of contracts whose audit verdict changes, on 175 Solidity contracts (143 vulnerable from SmartBugs-curated, 32 patched SWC Registry samples) with a fixed JSON-verdict prompt and a shared parser. Qwen2.5-Coder-0.5B, Llama-3.2-1B and Qwen3-0.6B were run at F16, Q8_0 and Q4_K_M with greedy decoding; an F16 repeat reproduced all outputs byte for byte. At full precision all three flag nearly every contract as vulnerable, so the stated vulnerability class carries the signal, and it churns heavily. For Qwen2.5-Coder, binary accuracy is identical at all precisions while 18.3% (Q8_0) and 45.7% (Q4_K_M) of class verdicts change, with a net loss of correct classes at Q4_K_M (14 vs. 1, McNemar p = 0.001) driven by self-contradictory outputs (20.0%→57.1%). Llama's Q4_K_M changes 37.1% of classes while class accuracy slightly improves; its three apparent misses stem from format degradation and vanish under a first-object parser. Churn is model-dependent: Q8_0 changed 1.1–18.3% of classes and no binary verdict, whereas Qwen3 at Q4_K_M lost 15 vulnerable-contract detections (p = 0.0005; accuracy 0.811→0.731), partly by ignoring its /no_think instruction. Across 21 API endpoints recall is 0.951–1.000 but specificity only 0.400–0.813; router caching makes identical repeats uninformative, whitespace-perturbed repeats bound run-to-run binary churn at 1.2–1.8%, and one model name served by two differently configured providers disagrees on 14 of 172 binary verdicts (p = 0.013). Paired verdict agreement should be reported alongside accuracy. This record contains the preprint PDF and the complete release package (code, dataset builder, contract set, raw model outputs for all local and API runs, results tables and LaTeX source). Code: https://github.com/horizonbymuneeb/smart-contract-verdict-churn

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-08
DOI
https://doi.org/10.5281/zenodo.23241639
Primary Topic
Blockchain Technology Applications and Security
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Same Accuracy, Different Verdicts: Quantization-Induced Verdict Churn in Small LLM Smart-Contract Auditors

Muneeb Anjum
Zenodo (CERN European Organization for Nuclear Research)
Blockchain Technology Applications and Security
preprint

Same Accuracy, Different Verdicts: Quantization-Induced Verdict Churn in Small LLM Smart-Contract Auditors

Muneeb Anjum
preprint en

Abstract

Small open LLMs are increasingly run as smart-contract auditors in quantized GGUF form, and large models are reached through API aggregators that do not disclose how they are served; both practices are usually justified by aggregate accuracy. We measure verdict churn, the share of contracts whose audit verdict changes, on 175 Solidity contracts (143 vulnerable from SmartBugs-curated, 32 patched SWC Registry samples) with a fixed JSON-verdict prompt and a shared parser. Qwen2.5-Coder-0.5B, Llama-3.2-1B and Qwen3-0.6B were run at F16, Q8_0 and Q4_K_M with greedy decoding; an F16 repeat reproduced all outputs byte for byte. At full precision all three flag nearly every contract as vulnerable, so the stated vulnerability class carries the signal, and it churns heavily. For Qwen2.5-Coder, binary accuracy is identical at all precisions while 18.3% (Q8_0) and 45.7% (Q4_K_M) of class verdicts change, with a net loss of correct classes at Q4_K_M (14 vs. 1, McNemar p = 0.001) driven by self-contradictory outputs (20.0%→57.1%). Llama's Q4_K_M changes 37.1% of classes while class accuracy slightly improves; its three apparent misses stem from format degradation and vanish under a first-object parser. Churn is model-dependent: Q8_0 changed 1.1–18.3% of classes and no binary verdict, whereas Qwen3 at Q4_K_M lost 15 vulnerable-contract detections (p = 0.0005; accuracy 0.811→0.731), partly by ignoring its /no_think instruction. Across 21 API endpoints recall is 0.951–1.000 but specificity only 0.400–0.813; router caching makes identical repeats uninformative, whitespace-perturbed repeats bound run-to-run binary churn at 1.2–1.8%, and one model name served by two differently configured providers disagrees on 14 of 172 binary verdicts (p = 0.013). Paired verdict agreement should be reported alongside accuracy. This record contains the preprint PDF and the complete release package (code, dataset builder, contract set, raw model outputs for all local and API runs, results tables and LaTeX source). Code: https://github.com/horizonbymuneeb/smart-contract-verdict-churn

Zenodo (CERN European Organization for Nuclear Research)
Blockchain Technology Applications and Security
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Same Accuracy, Different Verdicts: Quantization-Induced Verdict Churn in Small LLM Smart-Contract Auditors — Muneeb Anjum · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS