A research team from the University of Illinois Urbana-Champaign and IBM Research has proposed a way for large-language-model services to rebalance running GPU capacity before reaching for more hardware. The system, called FluidPD, moves work between the two main phases of inference and can also switch an active worker from one phase to the other without reloading the model. In experiments derived from Microsoft Azure workload traces, the researchers report gains of as much as 94.6 percentage points in the share of requests meeting both response-latency targets.
The paper addresses a growing infrastructure problem created by prefill-decode disaggregation. In that architecture, one pool of workers processes the user’s input prompt, or prefill, while another produces the answer token by token, or decode. The split lets operators optimize each phase separately, but it also fixes capacity into two specialized pools. A configuration that suits long conversational outputs may become a poor fit when traffic shifts toward code requests with much longer prompts.

FluidPD responds at two timescales. Its FluidToken mechanism handles brief prefill bursts by sending a bounded suffix of selected prompt work to a paired decode worker that has spare execution time and memory. Its FluidRole mechanism addresses sustained imbalance by draining a worker safely and changing its role in place. The model weights and runtime remain loaded, avoiding the startup and model-loading delay of adding a fresh worker.
Two predictive signals guide those decisions. A Prefill Pressure Index estimates whether queued prompts can still finish within their time-to-first-token budgets. A Decode Pressure Index tracks active decode load and available key-value-cache capacity. FluidPD permits work stealing only when the decode side has enough headroom, and it declines to move a worker when both phases are under sustained pressure. In that case, the paper says the service needs admission control or conventional autoscaling because the deployment lacks enough total capacity.
The researchers implemented FluidPD on SGLang 0.5.4 and tested Llama 3.1 8B, Qwen3 14B and Qwen3 30B-A3B. Their two replay workloads were constructed from separate weeklong Azure code and conversation traces: one mixed 10,000-request workload captured short bursts, while a staged 30,000-request workload created a sustained shift from conversation traffic to code traffic and back. The experiments ran on an AWS server with eight Nvidia A100 80GB GPUs, one GPU per serving instance.

Against a static SGLang deployment using the same eight-GPU allocation, FluidPD improved overall service-level-objective attainment by up to 94.6 percentage points on the staged workload and up to 74.3 points on the mixed workload. In a separate six-GPU comparison at 20 requests per second, FluidPD reached 97.6% attainment on the mixed workload versus 9.3% for Nvidia Dynamo’s autoscaling setup. On the staged workload, the figures were 92.6% and 5.2%, respectively, even though Dynamo could grow to eight GPUs.
The paper attributes that gap largely to reaction time. In one test, Dynamo made a scale-out decision about five seconds after a burst began, but the added worker did not become ready for roughly 200 seconds. Kubernetes scheduling and resource provisioning took about half that delay, with runtime initialization, model loading and CUDA graph capture accounting for the rest. FluidPD instead reused capacity already online. Its own control calculation took less than 0.50 milliseconds per iteration; measured role changes took about one second from prefill to decode and up to 12 seconds in the other direction.
The findings are promising but bounded. This is an arXiv preprint, and the evaluation used controlled replays rather than a live customer deployment. FluidPD’s latency predictors were calibrated offline for each tested model and hardware setup, and the authors did not compare directly with several related systems whose public implementations were unavailable. At the highest offered loads, using decode capacity for prompt work also caused some degradation in per-output-token latency. The central result is therefore not that autoscaling is obsolete, but that many short-lived or phase-specific bottlenecks may be solved faster by rearranging existing GPU work before adding machines.

Comments
Loading comments…