SafeBits: Safety-Constrained Global Bit Budgeting for Error-Feedback Gradient Quantization

Communication cost is a persistent bottleneck in distributed deep learning, especially when model scale causes gradient synchronization to dominate iteration time. Gradient quantization reduces this cost, but under tight global budgets it can also produce severely distorted updates whose directions deviate from the intended optimization signal, leading to instabilities that classical error-feedback (EF) analyses do not explicitly rule out. We introduce Safety-Constrained Global Budgeting with Error Feedback (SCGB-EF), a framework that couples per-tensor safety screening with global bit allocation. At each iteration, every trainable tensor selects a quantization action from a discrete menu under a hard communication budget, while a lightweight screen discards candidates whose normalized distortion exceeds a threshold or whose cosine similarity with the EF-corrected update falls below a prescribed floor. The resulting admissible sets define an explicit feasibility boundary: if the budget is insufficient to support a safe assignment, SCGB-EF falls back to the minimum-cost admissible configuration and logs the budget violation. We show that this design preserves the standard nonconvex convergence guarantees of EF, with an explicit contraction parameter \(\alpha = \bar{d}\) induced directly by the distortion threshold. Experiments on CIFAR-10 with ResNet-18 and ImageNet-100 with ResNet-50 show that SCGB-EF improves robustness and yields competitive accuracy, while incurring communication trade-offs relative to fixed-bit and non-safety-aware baselines, while eliminating training collapses in tight-budget regimes where unscreened methods fail. On GPT-2 / WikiText-103, SCGB-EF achieves perplexity within 5.0% of full-precision training at 7.2 \(\times\) compression and outperforms the closest fixed-bit baseline at comparable communication cost by 17 perplexity points. All reported communication reductions and compression ratios quantify modeled transmitted bit volume, not measured end-to-end network wall-clock speedups; realizing physical latency improvements requires integration with packed quantized collective communication kernels.

Authors

Institutions

Publication Details

Journal
International Journal of Computational Intelligence Systems
Published
2026-09-30
DOI
https://doi.org/10.1007/s44196-026-01626-z
Primary Topic
Advanced Neural Network Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

SafeBits: Safety-Constrained Global Bit Budgeting for Error-Feedback Gradient Quantization

Fahed Alkhabbas, Sadi Alawadi, Lovepreet Singh
International Journal of Computational Intelligence Systems
Advanced Neural Network Applications
article

SafeBits: Safety-Constrained Global Bit Budgeting for Error-Feedback Gradient Quantization

Fahed Alkhabbas, Sadi Alawadi, Lovepreet Singh
article en

Abstract

Communication cost is a persistent bottleneck in distributed deep learning, especially when model scale causes gradient synchronization to dominate iteration time. Gradient quantization reduces this cost, but under tight global budgets it can also produce severely distorted updates whose directions deviate from the intended optimization signal, leading to instabilities that classical error-feedback (EF) analyses do not explicitly rule out. We introduce Safety-Constrained Global Budgeting with Error Feedback (SCGB-EF), a framework that couples per-tensor safety screening with global bit allocation. At each iteration, every trainable tensor selects a quantization action from a discrete menu under a hard communication budget, while a lightweight screen discards candidates whose normalized distortion exceeds a threshold or whose cosine similarity with the EF-corrected update falls below a prescribed floor. The resulting admissible sets define an explicit feasibility boundary: if the budget is insufficient to support a safe assignment, SCGB-EF falls back to the minimum-cost admissible configuration and logs the budget violation. We show that this design preserves the standard nonconvex convergence guarantees of EF, with an explicit contraction parameter \(\alpha = \bar{d}\) induced directly by the distortion threshold. Experiments on CIFAR-10 with ResNet-18 and ImageNet-100 with ResNet-50 show that SCGB-EF improves robustness and yields competitive accuracy, while incurring communication trade-offs relative to fixed-bit and non-safety-aware baselines, while eliminating training collapses in tight-budget regimes where unscreened methods fail. On GPT-2 / WikiText-103, SCGB-EF achieves perplexity within 5.0% of full-precision training at 7.2 \(\times\) compression and outperforms the closest fixed-bit baseline at comparable communication cost by 17 perplexity points. All reported communication reductions and compression ratios quantify modeled transmitted bit volume, not measured end-to-end network wall-clock speedups; realizing physical latency improvements requires integration with packed quantized collective communication kernels.

International Journal of Computational Intelligence Systems
Malmö University (SE), United Arab Emirates University (AE), Indian Institute of Technology Madras (IN)
Partnerships for the goals
Openalex Percentile: Top 14%
Advanced Neural Network Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.