Latest Research in Accelerator Microarchitecture Optimization
245 research papers · 0.0 average citations · 2026 median publication year
Top Research Topics in Accelerator Microarchitecture Optimization
- Distributed, Parallel, and Cluster Computing — 46 papers
- Hardware Architecture — 40 papers
- Parallel Computing and Optimization Techniques — 24 papers
- Machine Learning — 10 papers
- Performance — 7 papers
- Programming Languages — 7 papers
- Scientific Computing and Data Management — 6 papers
- Information Theory — 6 papers
- Mathematical Software — 5 papers
- Networking and Internet Architecture — 5 papers
Highest-Cited Papers
- AI-coupled HPC Workflow Applications, Middleware and Performance (11 citations)
- VoltGrid: Microsecond Collective Interposition for Transient dI/dt Mitigation in Multi-Accelerator Training Clusters
- Twelve quick tips for designing AI-driven HPC workflows
- VoltGrid: Microsecond Collective Interposition for Transient dI/dt Mitigation in Multi-Accelerator Training Clusters
- The Price of the Bottleneck: A Pre-Registered, Multi-Domain Deployment Characterization of a Deterministic Edge Decision Token
- Machine-Learned Mismatch and Task Preservation Beliefs in CoSMA DAI for Common Knowledge Aware Semantic Alignment
- Comparative Analysis of Classical Machine Learning Vs. Deep Learning for Resource-Constrained Edge Devices
- PipeGNN: A Bandwidth-Efficient GNN Accelerator with Node-Level Pipelined Push Execution
- Scaling Fourier-Based Sparse Matrix Analysis on GPUs
- Chiral Topological Edge Routing for Thermal Mitigation and Bisection Decongestion in 2D Mesh Networks-on-Chip
- SemABR: Measuring Video Semantic Fidelity with Multimodal LLMs for Adaptive Bitrate Streaming
- Zero-I/O Fault Recovery for Sharded Deep Learning via Dynamic Framework Dependency Rebinding
- Chiral Topological Edge Routing for Thermal Mitigation and Bisection Decongestion in 2D Mesh Networks-on-Chip
- cosmokdtree: a flexible OpenMP-parallelized k-d tree for computational astrophysics applications
- Ozaki Scheme II: A GEMM-oriented emulation of floating-point matrix multiplication using an integer modular technique
- Quantifying the Effect of HCLs on a Fixed-Microarchitecture MXFP4 Accelerator
- MeshKV: A Network-on-Chip KV Cache Fabric for Scalable Transformer Decoding Accelerators
- Automated Instruction Encoding Synthesis for Modern GPU ISA Compression
- Accelerating Stateful Network Applications with Performance Prediction on SoC SmartNICs
- COMPASS-ABS: Reducing Fragmentation in Shared GPU Clusters for Deep Learning Training Workloads
Sub-Regions
- Hardware Architecture — 56 papers
- Information Theory — 54 papers
- Machine Learning — 44 papers
- Distributed, Parallel, and Cluster Computing — 30 papers
- Parallel Computing and Optimization Techniques — 22 papers
- Software Testing and Debugging Techniques — 8 papers