Latest Research in Accelerator Microarchitecture Optimization

245 research papers · 0.0 average citations · 2026 median publication year

Top Research Topics in Accelerator Microarchitecture Optimization

Highest-Cited Papers

  1. AI-coupled HPC Workflow Applications, Middleware and Performance (11 citations)
  2. VoltGrid: Microsecond Collective Interposition for Transient dI/dt Mitigation in Multi-Accelerator Training Clusters
  3. Twelve quick tips for designing AI-driven HPC workflows
  4. VoltGrid: Microsecond Collective Interposition for Transient dI/dt Mitigation in Multi-Accelerator Training Clusters
  5. The Price of the Bottleneck: A Pre-Registered, Multi-Domain Deployment Characterization of a Deterministic Edge Decision Token
  6. Machine-Learned Mismatch and Task Preservation Beliefs in CoSMA DAI for Common Knowledge Aware Semantic Alignment
  7. Comparative Analysis of Classical Machine Learning Vs. Deep Learning for Resource-Constrained Edge Devices
  8. PipeGNN: A Bandwidth-Efficient GNN Accelerator with Node-Level Pipelined Push Execution
  9. Scaling Fourier-Based Sparse Matrix Analysis on GPUs
  10. Chiral Topological Edge Routing for Thermal Mitigation and Bisection Decongestion in 2D Mesh Networks-on-Chip
  11. SemABR: Measuring Video Semantic Fidelity with Multimodal LLMs for Adaptive Bitrate Streaming
  12. Zero-I/O Fault Recovery for Sharded Deep Learning via Dynamic Framework Dependency Rebinding
  13. Chiral Topological Edge Routing for Thermal Mitigation and Bisection Decongestion in 2D Mesh Networks-on-Chip
  14. cosmokdtree: a flexible OpenMP-parallelized k-d tree for computational astrophysics applications
  15. Ozaki Scheme II: A GEMM-oriented emulation of floating-point matrix multiplication using an integer modular technique
  16. Quantifying the Effect of HCLs on a Fixed-Microarchitecture MXFP4 Accelerator
  17. MeshKV: A Network-on-Chip KV Cache Fabric for Scalable Transformer Decoding Accelerators
  18. Automated Instruction Encoding Synthesis for Modern GPU ISA Compression
  19. Accelerating Stateful Network Applications with Performance Prediction on SoC SmartNICs
  20. COMPASS-ABS: Reducing Fragmentation in Shared GPU Clusters for Deep Learning Training Workloads

Sub-Regions

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
L2 Region - - 2026 Sep Q3

Accelerator Microarchitecture Optimization

245 papers

Top Topics (10)

Distributed, Parallel, and Cluster Computing46
Hardware Architecture40
Parallel Computing and Optimization Techniques24
Machine Learning10
Performance7
Programming Languages7
Scientific Computing and Data Management6
Information Theory6
Mathematical Software5
Networking and Internet Architecture5

Top Publications (20)

1.AI-coupled HPC Workflow Applications, Middleware and Performance11c2.VoltGrid: Microsecond Collective Interposition for Transient dI/dt Mitigation in Multi-Accelerator Training Clusters3.Twelve quick tips for designing AI-driven HPC workflows4.VoltGrid: Microsecond Collective Interposition for Transient dI/dt Mitigation in Multi-Accelerator Training Clusters5.The Price of the Bottleneck: A Pre-Registered, Multi-Domain Deployment Characterization of a Deterministic Edge Decision Token6.Machine-Learned Mismatch and Task Preservation Beliefs in CoSMA DAI for Common Knowledge Aware Semantic Alignment7.Comparative Analysis of Classical Machine Learning Vs. Deep Learning for Resource-Constrained Edge Devices8.PipeGNN: A Bandwidth-Efficient GNN Accelerator with Node-Level Pipelined Push Execution9.Scaling Fourier-Based Sparse Matrix Analysis on GPUs10.Chiral Topological Edge Routing for Thermal Mitigation and Bisection Decongestion in 2D Mesh Networks-on-Chip11.SemABR: Measuring Video Semantic Fidelity with Multimodal LLMs for Adaptive Bitrate Streaming12.Zero-I/O Fault Recovery for Sharded Deep Learning via Dynamic Framework Dependency Rebinding13.Chiral Topological Edge Routing for Thermal Mitigation and Bisection Decongestion in 2D Mesh Networks-on-Chip14.cosmokdtree: a flexible OpenMP-parallelized k-d tree for computational astrophysics applications15.Ozaki Scheme II: A GEMM-oriented emulation of floating-point matrix multiplication using an integer modular technique16.Quantifying the Effect of HCLs on a Fixed-Microarchitecture MXFP4 Accelerator17.MeshKV: A Network-on-Chip KV Cache Fabric for Scalable Transformer Decoding Accelerators18.Automated Instruction Encoding Synthesis for Modern GPU ISA Compression19.Accelerating Stateful Network Applications with Performance Prediction on SoC SmartNICs20.COMPASS-ABS: Reducing Fragmentation in Shared GPU Clusters for Deep Learning Training Workloads

Sub-Regions (6)

AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.