Model Gateway
A model gateway is an intermediary service through which an application sends requests to one or more model providers. It exposes a stable application-facing endpoint while centralizing selected operational functions such as provider authentication, request translation, usage and cost tracking, rate limits, retries, fallbacks, caching, and observability. Implementations vary: a gateway may preserve provider-native payloads, present a common schema, or support both. The term describes the model-inference traffic layer, not every component of an AI platform.
Origin and context
Portkey publicly named an AI Gateway in August 2023 and described it as a layer between LLM applications and their providers, with logging, semantic caching, load balancing, and rate-limit handling. Cloudflare followed in September with an AI Gateway positioned between applications and AI APIs, including logging, caching, limiting, retries, and fallback endpoints. Vercel's 2025 general-availability release independently used the category for a unified API with authentication, usage tracking, and provider failover.
Why it matters
Without a shared traffic layer, each application may duplicate provider credentials, error handling, spend controls, telemetry, and switching logic. A gateway can make those concerns consistent across teams and can reduce application changes when a provider or model changes. It also creates one place to observe model calls and enforce approved routing choices. Those benefits are operational rather than semantic: the gateway does not itself make a model's answer correct, safe, or suitable for a task.
Example
A product team can point its chat service at one gateway endpoint, authorize the project with a scoped key, set a budget and rate limit, and record latency, token use, and errors. If the preferred provider is unavailable, the gateway may retry an approved deployment or invoke a configured fallback. The team should test the complete fallback path, because a request that another provider accepts may still differ in supported features, model behavior, latency, or output format. Routing policy and observability remain part of the application design.
How it differs
LLMOps
LLMOps is broader than the gateway layer. Portkey's 2023 description listed its AI Gateway alongside observability, prompt management, experimentation and evals, and security and compliance as separate parts of an LLMOps platform. A team can therefore use a model gateway without treating it as the whole operating lifecycle.
Prompt Caching
Prompt caching reuses recently processed input tokens to reduce repeated work, cost, or latency. A model gateway mediates provider traffic and may offer a cache, but caching is optional and can also be implemented by a model provider without a separate gateway. The two capabilities should be configured and evaluated independently.
Maturity and evidence
Maturity is rated 3. Independent implementations from Portkey, Cloudflare, and Vercel converge on a recognizable gateway boundary and on recurring capabilities such as unified access, tracking, retries, and fallback. The category is established enough for architecture decisions, but interfaces, feature coverage, policy semantics, and provider compatibility are not standardized across products.
Limits and open questions
Placing a gateway in the request path adds a dependency whose availability, latency, configuration, and credentials must be operated and tested. Central logging, caching, and cost records may process sensitive prompts or outputs, so access, retention, redaction, and cache partitioning need explicit policies. Fallback must not assume that different models are behaviorally interchangeable. Teams should allowlist providers, test tool and structured-output compatibility, bound retries, and keep deterministic authorization outside probabilistic model decisions. A gateway can carry policy controls, but the label alone is not evidence that adequate controls exist.
Related terms
References
- Announcing AI Gateway: making AI applications more observable, reliable, and scalableCloudflare · 2023-09-27 · class A
- Announcing $3M Seed Round to Bring LLMs to ProductionPortkey · 2023-08-23 · class B
- AI Gateway: Production-ready reliability for your AI appsVercel · 2025-08-21 · class A
- Prompt Caching in the APIOpenAI · 2024-10-01 · class A
Last updated: 2026-09-04
This term is also covered in the Skills Atlas as llm api gateway skill.
This term is also covered in the Skills Atlas as llm api integration skill.
This term is also covered in the Skills Atlas as llm observability skill.