Latest Research in Distributed, Parallel, and Cluster Computing
57 research papers · 2026 median publication year
Top Research Topics in Distributed, Parallel, and Cluster Computing
- Machine Learning — 15 papers
- Distributed, Parallel, and Cluster Computing — 14 papers
- Hardware Architecture — 4 papers
- Computer Vision and Pattern Recognition — 3 papers
- Machine Learning — 2 papers
- Domain Adaptation and Few-Shot Learning — 2 papers
- Artificial Intelligence — 2 papers
- Adversarial Robustness in Machine Learning — 2 papers
- Computation and Language — 2 papers
- Caching and Content Delivery — 2 papers
Highest-Cited Papers
- The Output-Space Hypothesis: Enumerative Equivalence Checking for Tensor Programs
- PixelFlow: Token-Level Workload Management for Efficient Distributed DiT Serving
- MiX: Micro-Inverted-Scaling for End-to-End Low-Bit Vision-Language Model Acceleration
- Zero-I/O Fault Recovery for Sharded Deep Learning via Dynamic Framework Dependency Rebinding
- Finite-Sample Unbiased Variance of MMD under Unbalanced Sampling: Exact Estimation and Quasi-Linear Computation
- Optimal Post-processing of Synthetic Data for Pearson Correlation Matching
- s-MDM: Generative Virtualization of Multi-Device Hardware Variations for Portable DL-SCA
- RayOrch: Programming and Executing Lineage-Controlled Multi-Grain Dataflows for Foundation-Model Data Preparation
- Epsilon-Nash Equilibria in History-Dependent SA-MDPs
- NCCLZ: Compression-Enabled GPU Collectives with Decoupled Quantization and Entropy Coding
- TAST: Task-Aware Sparse Topology Learning for Multi-Task Recommendation
- FairLRF: Achieving Fairness through Sparse Low Rank Factorization
- Online Allocation using Few Samples
- FairCompressAgent: An Agentic Framework for Fairness-Aware Model Compression for FPGA Deployment
- DeepSeek-V4-Flash on AMD gfx90a: Correctness Recovery and Inference Performance Engineering
- The World Model Hardware Accelerator
- DMA-Latte: Expanding the Reach of DMA Offloads to Latency-bound ML Communication
- Partition-Aware Scheduling for Mobile Heterogeneous Inference Co-Execution
- Affinity-Aware Sharding for Delayed Tensor Parallelism
- Distributed Mechanistic Interpretability at Scale: Activation Streaming, Split-Layer Inference, and Distributed Sparse Autoencoder Training