Empirical Tail-Latency Characterization of GPU-Accelerated 5G NR LDPC Decoding
GPU LDPC latency comparisons across implementations are challenging due to variations in decoder schedules, batch granularity, and timing boundaries. Batch completion was characterized for FP32-layered and flooding decoders on a GPU, considering block acquisition, matched outcomes, and three timing boundaries. The schedule study included 600,000 observations. Flooding exhibited higher P50 and P99 in 15 full-boundary matched-cap comparisons, while J99 and R99 did not display uniform schedule ordering. For words satisfying the syndrome under both schedules, layered decoding achieved first satisfaction earlier in every estimable 2- and 4-dB cell. Layered endpoint levels varied among full, decode-only, and transfer-only boundaries. Over six sessions, full-boundary J99 increased by 234.7 μs from batch 64 to 128 (pointwise descriptive 95% percentile interval, 196.1–281.5 μs), which contrasted with two other designs. These results pertain to a single dynamically clocked Windows Subsystem for Linux 2 (WSL2) GPU and do not establish fixed-clock or cross-platform behavior, radio-interface latency, causal attribution, or worst-case bounds.
Authors
- Sooyoung Jang (ORCID: https://orcid.org/0000-0002-6931-9592)
- Eunkyung Kim (ORCID: https://orcid.org/0000-0003-3558-7086)
Institutions
- Hanbat National University (KR)
Publication Details
- Journal
- Telecom
- Published
- 2026-09-15
- DOI
- https://doi.org/10.3390/telecom7050120
- Primary Topic
- Error Correcting Code Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00