Imagine an industrial kitchen where one chef is tasked with simultaneously doing two completely opposite jobs: chopping 500 pounds of onions at maximum physical speed (high-intensity compute burst), and gently stirring a delicate sauce one drop every three seconds (slow, sequential monitoring). If you force both tasks onto the same chef at the exact same moment, the chef chops irregularly and burns the sauce. The solution is obvious: separate the kitchen into a dedicated prep station and a dedicated finishing station. This is Disaggregated Inference.
The Two Conflicting Personalities of LLM Inference
Every language model request consists of two mathematically distinct computational phases:
- The Prefill Phase (Prompt Processing): The model ingests a 10,000-token prompt and computes self-attention across all tokens simultaneously. This phase is Compute-Bound (FLOPS-heavy). It saturates GPU Tensor Cores at 90%+ utilization and finishes in a single massive parallel burst.
- The Decode Phase (Token Generation): The model generates output text one token at a time. This phase is Memory-Bandwidth Bound. It uses almost zero arithmetic compute, but requires streaming the entire 140GB model weights across memory buses for every single output token.
[Colocated Inference: Mutual Resource Interference]
Incoming Request ──► [Single GPU: Forced to do both Prefill & Decode simultaneously]
├── Prefill bursts cause microsecond stuttering in active streaming users!
└── Decode memory transfers stall incoming prefill processing.
[Disaggregated Inference: Split-Phase Architecture]
Incoming Prompt ──► [Dedicated PREFILL Pool (Optimized for TeraFLOPS Compute)]
│
▼ (Ultra-Fast RDMA Transfer of Key-Value Cache)
[Dedicated DECODE Pool (Optimized for High Bandwidth Memory)]
│
▼
Smooth, Uninterrupted Token Streaming (Zero Jitter!)
The Interference Problem in Colocated Serving
When an inference server mixes prefill and decode tasks on the same GPU, catastrophic resource contention occurs: an incoming user submitting a large document causes a massive compute spike that temporarily halts the streaming token output of ten other users already in the middle of reading their answers.
The Disaggregated Solution: Dedicated Pools & RDMA Cache Transfer
Pioneered by architectures like DistServe and Mooncake, Disaggregated Serving separates the cluster into two specialized hardware pools:
- Prefill Pool: Packed with high-TFLOPS compute GPUs that process prompts at breakneck speed.
- Decode Pool: Packed with high-memory-bandwidth GPUs optimized for maximum concurrent streaming throughput.
- RDMA KV-Cache Transfer: As soon as the prefill node completes the prompt attention matrix, it transmits the pre-computed Key-Value cache across ultra-fast RoCE / InfiniBand RDMA directly into the memory of a decode worker in under 5 milliseconds.
Engineering Takeaway
Never mix compute-bound workloads with memory-bound streaming workloads on the same silicon. By disaggregating prefill and decode execution paths, you eliminate latency jitter and maximize the operational efficiency of your GPU fleet.