Model Router

AI Infrastructure Intermediate

A model router is the component, usually inside an AI gateway, that decides which specific model handles a given request, based on rules like cost, latency, task type, or current load.

In simple terms

Not every request needs the most expensive, most capable model. A router can send a simple classification task to a small, cheap, fast model and reserve the largest model for genuinely hard requests, automatically, based on rules set in advance.

Why it matters

It's a direct lever on cost and speed at scale: routing the right request to the right model, instead of sending everything to one model regardless of difficulty, is one of the highest-leverage optimizations available in a production AI system.

How it works

Routing rules can be simple (a fixed mapping from task type to model) or dynamic (choosing based on real-time latency, current cost budget, or a classifier that estimates task difficulty first). Some routers also split traffic for A/B testing or gradually shift volume to a newly deployed model.

Where it fits

Incoming request → Model router → Routing decision (cost, latency, task type) → Selected model → Response

Production impact

A well-tuned router can cut average cost significantly by reserving expensive models for genuinely hard requests, this site's own model-comparison data (which model is fastest, cheapest, or most accurate per task type) is exactly the kind of input a router's rules would be built on.

On howfastai.com

Choosing the right model for a task, the whole premise of this site's Which Model pages, is precisely the decision a model router automates at request time instead of a human making it once in advance.

Learn this first

Related terms

Learn next

← All terms · Knowledge map →