Model serving is the infrastructure layer that takes a trained model and makes it available to answer real requests, handling batching, queueing, and GPU allocation so many requests can be served efficiently at once.
A trained model by itself is just a large file of numbers. Model serving is everything that turns that file into something you can send a prompt to and get an answer back from, quickly, reliably, and for many users at the same time.
The same model, served well versus served naively, can differ by an order of magnitude in throughput and cost. Serving infrastructure is where most of the real engineering effort in production LLM systems actually goes.
A serving system loads model weights onto GPU memory, accepts incoming requests, and groups them into batches for efficient GPU use (continuous batching lets new requests join a batch mid-generation rather than waiting for a batch to fully finish). It manages the KV cache across all concurrent requests, schedules prefill and decode work, and streams tokens back as they're generated.
Serving efficiency directly sets how many requests a fixed number of GPUs can handle, and therefore the real per-request cost, independent of the model's own size or architecture.