An AI gateway is a layer that sits between an application and the various AI model providers it uses, handling routing, rate limiting, retries, and cost tracking through one consistent interface instead of the app calling each provider directly.
If an application calls three different model providers directly, it has to handle each one's quirks, auth, rate limits, error formats, itself, and duplicate that logic three times over. A gateway centralizes all of that: the app talks to one endpoint, and the gateway figures out which provider actually handles the request.
As products increasingly use multiple models for different tasks, or need to fail over from one provider to another during an outage, a gateway is what makes that manageable instead of a tangle of provider-specific code scattered through the application.
Requests come into the gateway in a normalized format; it applies rate limiting and quota checks, selects a target model or provider (by explicit routing rules, cost, or latency), forwards the request, and handles retries or fallback to a different provider on failure. It typically also logs usage for cost tracking and observability along the way.
A gateway adds a small amount of latency to every request, but usually pays for that in reliability (automatic fallback during a provider outage) and negotiating leverage (usage aggregated across providers is easier to optimize for cost).