Symmetry-Aware INT4 Quantized GNN Decoder: ASIC Synthesis and Low-Latency Surface Code Architecture

Real-time quantum error correction for superconducting processors requires decoding streaming syndrome data within microsecond cycle intervals (T_cycle ≈ 1.1 μs). Classical minimum-weight perfect matching executed on host processors suffers from communication and serial matching bottlenecks, creating an exponential decoding backlog that limits quantum execution. This paper presents a design automation and hardware-software co-design framework for real-time surface code decoding using symmetry-aware, low-bit quantized graph neural networks. Operating as a confidence-gated hardware pre-filter, the four-bit integer (INT4) decoder commits 73.4% to 94.0% of syndrome frames directly on silicon within deterministic sub-microsecond deadlines, routing only ambiguous degenerate frames to classical matching while preserving full logical fidelity. To prevent arithmetic collapse in fixed-point datapaths without increasing word length, lattice dihedral equivariance is incorporated as an algebraic variance regularizer that suppresses activation outliers. Automated clique projection converts irregular detector hypergraphs into bounded-degree topologies, enabling an initiation interval of one cycle on systolic pipelines. Evaluated across physical AMD Kintex UltraScale+ FPGA hardware telemetry and sign-off post-route SkyWater 130 nm standard-cell ASIC synthesis (room-temperature extractions at 25°C and 1.8 V evaluated analytically against a 1.5 W 4K cryostat cooling lift), the architecture delivers 106.2 to 396.8 ns execution latency with +0.18 ns positive static timing slack at 312.5 MHz clock frequency and 3.80 to 66.25 nJ energy per operation. Multi-server queueing analysis confirms queue stability under continuous 1 MHz syndrome streams, bounding average queue wait time well within physical qubit coherence limits. Note: This work has been submitted to the IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems (TCAD) for possible publication.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-26
DOI
https://doi.org/10.5281/zenodo.22978989
Primary Topic
Quantum Computing Algorithms and Architecture
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Symmetry-Aware INT4 Quantized GNN Decoder: ASIC Synthesis and Low-Latency Surface Code Architecture

Maheli Ahmed, Md. Nazmul, Musrat Jahan Gungun
Zenodo (CERN European Organization for Nuclear Research)
Quantum Computing Algorithms and Architecture
preprint

Symmetry-Aware INT4 Quantized GNN Decoder: ASIC Synthesis and Low-Latency Surface Code Architecture

Maheli Ahmed, Md. Nazmul, Musrat Jahan Gungun
preprint en

Abstract

Real-time quantum error correction for superconducting processors requires decoding streaming syndrome data within microsecond cycle intervals (T_cycle ≈ 1.1 μs). Classical minimum-weight perfect matching executed on host processors suffers from communication and serial matching bottlenecks, creating an exponential decoding backlog that limits quantum execution. This paper presents a design automation and hardware-software co-design framework for real-time surface code decoding using symmetry-aware, low-bit quantized graph neural networks. Operating as a confidence-gated hardware pre-filter, the four-bit integer (INT4) decoder commits 73.4% to 94.0% of syndrome frames directly on silicon within deterministic sub-microsecond deadlines, routing only ambiguous degenerate frames to classical matching while preserving full logical fidelity. To prevent arithmetic collapse in fixed-point datapaths without increasing word length, lattice dihedral equivariance is incorporated as an algebraic variance regularizer that suppresses activation outliers. Automated clique projection converts irregular detector hypergraphs into bounded-degree topologies, enabling an initiation interval of one cycle on systolic pipelines. Evaluated across physical AMD Kintex UltraScale+ FPGA hardware telemetry and sign-off post-route SkyWater 130 nm standard-cell ASIC synthesis (room-temperature extractions at 25°C and 1.8 V evaluated analytically against a 1.5 W 4K cryostat cooling lift), the architecture delivers 106.2 to 396.8 ns execution latency with +0.18 ns positive static timing slack at 312.5 MHz clock frequency and 3.80 to 66.25 nJ energy per operation. Multi-server queueing analysis confirms queue stability under continuous 1 MHz syndrome streams, bounding average queue wait time well within physical qubit coherence limits. Note: This work has been submitted to the IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems (TCAD) for possible publication.

Zenodo (CERN European Organization for Nuclear Research)
National University Bangladesh (BD)
Quantum Computing Algorithms and Architecture
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.