vLLM is an open-source inference engine built around paged attention, a technique that manages the KV cache in fixed-size blocks (like an operating system manages memory pages) to dramatically reduce wasted GPU memory.
Naive KV cache allocation reserves memory upfront for the longest possible response, wasting most of it on shorter ones. vLLM's paged attention instead allocates cache memory in small chunks as needed, the way an operating system pages memory to processes, so far less GPU memory sits idle and reserved.
It's one of the most widely adopted open-source inference engines specifically because paged attention lets a fixed set of GPUs serve substantially more concurrent requests than naive serving, without any change to the model itself.
Instead of one contiguous memory block per request, the KV cache is split into fixed-size pages that can be allocated non-contiguously and shared between requests where sequences overlap (like a shared system prompt). A block manager tracks which pages belong to which request, similar to page tables in an operating system, freeing pages as soon as a request finishes.
Higher achievable batch sizes from reduced memory waste translate directly into higher throughput and lower cost per request on the same hardware, without touching model quality at all.