vLLM

Inference & Serving Infrastructure

vLLM is an open-source inference engine built around paged attention, a technique that manages the KV cache in fixed-size blocks (like an operating system manages memory pages) to dramatically reduce wasted GPU memory.

In simple terms

Naive KV cache allocation reserves memory upfront for the longest possible response, wasting most of it on shorter ones. vLLM's paged attention instead allocates cache memory in small chunks as needed, the way an operating system pages memory to processes, so far less GPU memory sits idle and reserved.

Why it matters

It's one of the most widely adopted open-source inference engines specifically because paged attention lets a fixed set of GPUs serve substantially more concurrent requests than naive serving, without any change to the model itself.

How it works

Instead of one contiguous memory block per request, the KV cache is split into fixed-size pages that can be allocated non-contiguously and shared between requests where sequences overlap (like a shared system prompt). A block manager tracks which pages belong to which request, similar to page tables in an operating system, freeing pages as soon as a request finishes.

Where it fits

Model weights → vLLM engine (paged KV cache) → GPU: continuous batching → Served tokens

Production impact

Higher achievable batch sizes from reduced memory waste translate directly into higher throughput and lower cost per request on the same hardware, without touching model quality at all.

Learn this first

Related terms

Learn next

← All terms · Knowledge map →