Q-LATTE: QoS-Aware and Workload-Adaptive Local-Cloud Hybrid Storage for Multi-Tenant Cloud Instances
Cloud local storage has transitioned through three hardware generations: software polling (ESPRESSO), ASIC hardware offloading (DOPPIO), and ASIC/SoC co-design (RISTRETTO). The state-of-the-art hybrid architecture (LATTE, USENIX FAST '26) pairs fast local flash with Elastic Block Store (EBS) via an open-source Cloud Storage Acceleration Layer (CSAL) and a Linear-SVM classifier. However, in shared multi-tenant deployments, unthrottled write bursts trigger severe Quality-of-Service (QoS) degradation, inducing up to 3.8x tail-latency spikes at P99.9. Furthermore, single-tenant linear dispatchers fail to detect dynamic workload phase shifts. We propose Q-LATTE, an open-source, QoS-aware hybrid storage architecture featuring: (1) a lock-free Token-Bucket Fair-Share Admission Controller (T-BFAC) inside the SPDK polling loop, (2) a Two-Tier Phase-Aware Dispatcher (TPA) coupling sub-40ns spatial heuristics with online learning, and (3) an Asymmetric Eviction Policy prioritizing bulk cloud writebacks. Evaluated on a commodity NVMe testbed replicating Alibaba Cloud's production environment, Q-LATTE eliminates queue head-of-line blocking, reducing P99.9 write tail latency by 70.7% (345 us vs. 1,180 us) while boosting MySQL throughput by 24.6%. Open-source replication artifacts and dataset are available at: https://github.com/codenameyizzz/Reproducing-Cloud-Local-Storage-Evolution-LATTE-Hybrid-Tiering-USENIX-FAST-26
Authors
- Yizreel Schwartz Sipahutar (ORCID: https://orcid.org/0009-0002-1625-0971)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-06
- DOI
- https://doi.org/10.5281/zenodo.23182534
- Primary Topic
- Advanced Data Storage Technologies
- Type
- preprint