In the early days of multi-user operating systems, running multiple programs simultaneously was notoriously difficult because memory had to be allocated in contiguous physical blocks. The operating system revolution occurred when computer scientists invented Virtual Memory and Paging—breaking memory into small, uniform pages that could be scattered anywhere in physical RAM. In 2023, the exact same revolution transformed the world of large language model serving.
The Memory Fragmentation Crisis
When multiple users send prompts of unpredictable lengths to an AI server, managing their Key-Value (KV) caches is a nightmare. Early serving systems pre-allocated a contiguous chunk of GPU VRAM sized for the maximum possible sequence length (e.g., 8,192 tokens) for every incoming request.
If a user only asked a 100-token question, over 95% of that pre-allocated GPU VRAM sat completely empty and wasted, locked away from other users. As a result, expensive $30,000 GPUs ran out of memory after serving only a handful of concurrent users.
[Static Contiguous Memory: 70% VRAM Wasted in Fragmentation] Request 1 (100 tokens): [Data][====== 8,000 Tokens of Empty Wasted VRAM ======] Request 2 (300 tokens): [Data][====== 7,800 Tokens of Empty Wasted VRAM ======] [PagedAttention (vLLM): Zero Waste Virtual Memory] Physical GPU Memory Pages: [P1][P2][P3][P4][P5][P6][P7][P8] Request 1 Logical Tokens ──► Block Table ──► Maps dynamically to [P1, P4] Request 2 Logical Tokens ──► Block Table ──► Maps dynamically to [P2, P3, P7] (Memory waste drops below 4%, enabling 4x to 8x higher concurrent user throughput!)
Continuous Iteration-Level Batching
Traditional batching in deep learning grouped requests together and waited for the slowest, longest generation to finish before accepting new work. This caused fast 10-token queries to be held hostage by 2,000-token novel generation requests.
Continuous Batching (Cellular Batching): Instead of batching at the request level, modern engines batch at the iteration level. After every single forward pass (every token emitted), completed requests are evicted, new incoming requests are instantly slotted into the batch, and the GPU never wastes a single microsecond idling.
Engineering Takeaway
High-performance AI serving is fundamentally a systems engineering discipline. By combining virtual memory paging (PagedAttention) with dynamic continuous batching, modern engines deliver sub-second latencies and massive throughput on production workloads.