QoServe: QoS-Aware LLM Inference Serving
Building deadline-aware scheduling for shared LLM inference infrastructure.
Building deadline-aware scheduling for shared LLM inference infrastructure.
Systems work on transparent preemption, migration, and elasticity for large-scale AI workloads.
Published in ICACCE, 2019
Simulation framework for energy aware VM allocation in a cloud data center.
Recommended citation: Bhandia et al.
Download Paper
Published in arxiv, 2022
Pre-emptive and elastic scheduling of AI workloads at planet-scale.
Recommended citation: Shukla et al
Download Paper
Published in INCET, 2022
Task partitioning framework for heterogeneous systems.
Recommended citation: Yekbote et al.
Download Paper
Published in HiPC, 2022
Accelerating Key-value stores using Page Table Walkers.
Recommended citation: Anupindi et al.
Download Paper
Published in US Patent 12,498,935 B2, 2025
US Patent 12,498,935 B2. Transparent time-slicing of multiple training workers on a single accelerator device, using replica-aware memory placement to make elastic GPU sharing cheap.
Recommended citation: Sivathanu et al.
Download Paper
Published in ACM ASPLOS 2026, 2026
QoS-aware scheduling for LLM inference serving, co-scheduling latency-sensitive and tolerant workloads on shared infrastructure.
Recommended citation: Goel et al.
Download Paper