A Bandwidth-Balanced Adaptive Technique for Mesh-Based I/O Chiplets
As the semiconductor industry rapidly embraces chiplet-based architectures to satisfy AI-driven high-performance computing demands, the flexible, scalable I/O‑Die (IOD) chiplet offers a promising path toward solving the AI memory wall. However, the design of these IOD chiplets for flexibility and reuse across products dictates a fixed peripheral placement for all memory controllers and die-to-die interfaces, creating unique and largely unexplored challenges. This rigid layout causes severe, non-uniform congestion in the mesh-based Network-on-Chip (NoC), underutilizing its full cross-sectional bandwidth and forming a critical system-level performance bottleneck. To address this, we propose the Bandwidth Balanced Adaptive Technique (BBAT) that diversifies traffic across the mesh paths using a probabilistic routing hint updated online from measured congestion. Unlike prior adaptive techniques, BBAT’s core innovation is a feedback mechanism that leverages the existing high-performance cache-coherent transaction protocol, eliminating the need for custom signaling. To dynamically adapt to shifting traffic patterns, our technique employs simple, low-overhead hardware logic that processes both traditional and novel congestion metrics gathered from across the chiplet—a design optimized for high performance while maintaining power efficiency. We conduct an extensive evaluation using a cycle-accurate simulator across two distinct chiplet topologies—one modeled after a commercial, asymmetric IOD and another with a uniform interface layout. The results demonstrate that our technique achieves a peak performance improvement of 2.6 × over the Dimension-Ordered Routing (DOR) baseline in synthetic traffic and outperforms prior congestion-aware adaptive approaches by up to 14% under heavily congested synthetic patterns. Across a practical mix of synthetic traffic patterns evaluated on both topologies, BBAT consistently delivers a 1.1 × –1.8 × speedup relative to the deterministic DOR baseline. Critically, on memory-bound SPEC CPU2017 workloads, it achieves a 1.2 × –1.7 × speedup over DOR and 7–10% higher throughput than existing adaptive techniques, demonstrating effective load balancing and strong congestion responsiveness in high-performance chiplet architectures, while adding only 5.4% to total NoC area.
Authors
- Prasenjit Chakraborty (ORCID: https://orcid.org/0000-0001-8495-7405)
- Zehao Lu (ORCID: https://orcid.org/0009-0005-0028-9910)
Institutions
- American Rock Mechanics Association (US)
Publication Details
- Journal
- ACM Transactions on Architecture and Code Optimization
- Published
- 2026-10-06
- DOI
- https://doi.org/10.1145/3857352
- Primary Topic
- Interconnection Networks and Systems
- Type
- article
- Field-Weighted Citation Impact
- 0.00