What Is an LLM Gateway? How Routing, Caching, and Fallback Work

What Is an LLM Gateway? How Routing, Caching, and Fallback Work

TLDR: An LLM gateway is a middleware layer between your application and one or more model providers that handles authentication, rate limiting, fallback, cost tracking, and caching in one place. It differs from a bare proxy (no logic) and a router (picks a model, nothing else). This guide covers what LiteLLM, Portkey, and Kong AI Gateway actually implement, a tested fallback-and-cache core you can read end to end, and when adding one is worth the extra hop.

Gateway vs Router vs Proxy

The three terms get used interchangeably, but they name different amounts of logic sitting in front of a model call. A proxy forwards a request to one destination and adds little beyond authentication; it does not choose between providers. A router adds a decision: given a request, pick a model or provider by rule, such as sending coding prompts to a coding-tuned model. A gateway is the superset: routing plus policy enforcement (rate limits, budgets), observability (logging, spend tracking), and reliability features (fallback, retries, caching) in one layer that every call in your application goes through. In practice the three products people reach for in this space, LiteLLM, Portkey, and Kong's AI Gateway, all call themselves gateways because they all do more than forward requests.

What a Gateway Actually Adds, Feature by Feature

Stripped of marketing language, the features a gateway sells are consistently these five, and each has a real implementation worth knowing:

  • Provider fallback. LiteLLM's router lets you list backup deployments per model name; when a call fails, it retries against the next entry, and repeated failures put a deployment on a cooldown timer rather than retrying it immediately.
  • Rate limiting and spend tracking. LiteLLM's proxy tracks budgets per virtual key, team, or user and can enforce a hard cap. Kong's AI Gateway records usage into whatever logging plugin you already run in Kong, so cost data lands next to your existing API metrics instead of in a separate system.
  • Format normalization. Kong's AI Proxy plugin translates a single OpenAI-shaped request into the target provider's native format across OpenAI, Azure OpenAI, Bedrock, Anthropic, Gemini, Vertex AI, Cohere, Mistral, Hugging Face, DeepSeek, Ollama, and vLLM, then translates the response back. As of Kong Gateway 3.10, you can also set config.llm_format to skip translation and pass a provider's native SDK format straight through when you do not want normalization.
  • Caching. Storing a response so an identical or near-identical request does not re-hit the model. Exact-match caching only helps on literal duplicates; semantic caching, covered below, helps on paraphrases too.
  • Observability. Centralized logs of every request and response across providers, which is the feature every gateway leads with because it is the one you cannot easily bolt on after the fact once calls are scattered across your codebase.

A Fallback Config That Actually Runs

This is LiteLLM's own documented shape for a two-deployment fallback, not a paraphrase:

model_list:
  - model_name: primary-model
    litellm_params:
      model: azure/<your-deployment-name>
      api_base: <your-azure-endpoint>
      api_key: <your-azure-api-key>
      rpm: 6
  - model_name: backup-model
    litellm_params:
      model: azure/<your-backup-deployment>
      api_base: <your-backup-endpoint>
      api_key: <your-azure-api-key>
      rpm: 6
router_settings:
  fallbacks: [{"primary-model": ["backup-model"]}]

Two behaviors are easy to miss when you read this too quickly. First, rpm is a rate limit on that specific deployment, not a global budget, so a fallback entry with its own rpm gets its own quota rather than sharing the primary's. Second, LiteLLM's cooldowns apply to individual deployments: in a model group with several deployments, one that crosses the failure threshold is pulled from rotation for a cooldown window while its siblings keep serving. In this config each group has a single deployment, and LiteLLM's router source skips the 429 and high-failure-rate cooldowns for single-deployment groups by default, so a flaky primary is still tried first on every request and the fallback does the rerouting.

The Part No Vendor Page Shows You: A Gateway Core You Can Read End to End

Every product above wraps the same two mechanics in more configuration and more providers: try backends in priority order, and cache what succeeded. Here is that core with nothing hidden, tested against local fixture servers that stay healthy, return 500s, hang past the timeout, or drop the connection:

import hashlib, http.client, json, time, urllib.request


class GatewayError(Exception):
    def __init__(self, attempts):
        self.attempts = attempts
        super().__init__("all backends failed: %r" % (attempts,))


class Gateway:
    def __init__(self, backends, cache_ttl_seconds=300):
        self.backends = backends  # [(name, base_url, api_key), ...]
        self.cache_ttl = cache_ttl_seconds
        self._cache = {}
        self.stats = {"cache_hits": 0, "cache_misses": 0, "fallbacks": 0}

    def _cache_key(self, payload):
        basis = json.dumps(payload, sort_keys=True)  # every field can change the answer
        return hashlib.sha256(basis.encode("utf-8")).hexdigest()

    def chat(self, payload, timeout=5):
        key = self._cache_key(payload)
        cached = self._cache.get(key)
        if cached and cached[0] > time.time():
            self.stats["cache_hits"] += 1
            return dict(cached[1], _served_by="cache")
        self.stats["cache_misses"] += 1

        attempts = []
        for name, base_url, api_key in self.backends:
            try:
                resp = self._call(base_url, api_key, payload, timeout)
            except (OSError, ValueError, http.client.HTTPException) as e:
                # HTTP errors, timeouts, dropped connections, non-JSON bodies
                attempts.append((name, repr(e)))
                self.stats["fallbacks"] += 1
                continue
            self._cache[key] = (time.time() + self.cache_ttl, resp)
            return dict(resp, _served_by=name)
        raise GatewayError(attempts)

    def _call(self, base_url, api_key, payload, timeout):
        body = json.dumps(payload).encode("utf-8")
        req = urllib.request.Request(
            base_url.rstrip("/") + "/chat/completions", data=body, method="POST",
            headers={"Content-Type": "application/json",
                     "Authorization": "Bearer " + api_key})
        with urllib.request.urlopen(req, timeout=timeout) as resp:
            return json.loads(resp.read().decode("utf-8"))

Run against those fixtures, it passes the checks that matter for a fallback layer: it tries the primary first, falls back to the secondary exactly once when the primary fails (a 500, a hang past the timeout, or a dropped connection), serves the second identical call from cache without touching either backend again, treats a request with different message content or parameters as a cache miss, and raises a single GatewayError carrying every attempt when both backends are down. That last case matters in practice: swallowing the failure and returning nothing is worse than a clear exception naming every backend that was tried.

Caching: What the Published Numbers Actually Support

"Caching cuts your API costs" is true, but how much depends entirely on how repetitive your traffic is, and the honest number range comes from published, workload-specific results rather than a single flat figure. The original GPTCache paper reports a 2 to 10 times speedup on a cache hit, which is a latency result, not a cost-savings percentage. A later semantic-caching study first posted to arXiv in November 2024 measured the cost-relevant number directly: across four query categories, cache hit rates ran from 61.6% to 68.8%, meaning that fraction of API calls was avoided entirely. At the similarity threshold the authors used, a GPT-4o mini judge rated 92.5% to 97.3% of cache hits as accurate answers, depending on category; the abstract's "exceeding 97%" holds for only one of the four categories in the paper's own results table. Both results are workload-dependent: a chatbot answering the same handful of FAQ-style questions will land near the high end, and an application where every prompt is materially different will see close to nothing. Treat any caching claim you cannot trace to a measured workload the way you would treat an unsourced benchmark: as a best case, not a guarantee.

Gateways, Web Search, and Tool Calls

A gateway's routing and policy layer is not limited to model calls. When an agent behind the gateway needs to call a web search API as a tool, the same rate limiting, budget tracking, and fallback logic that applies to model calls can wrap that request too, since from the gateway's point of view it is just another outbound HTTP call with a cost and a failure mode. The You.com Web Search API fits this pattern over its REST endpoint, and its MCP server can be registered as a tool provider in gateway architectures that speak MCP, keeping search calls behind the same access point as model calls rather than as a separate, unmonitored dependency.

Three Tools, Compared on What They Actually Ship

ToolLicense modelCoverageNotable, sourced detail
LiteLLMOpen-source proxy core100+ LLM providers behind one interfaceOrder-based fallback with per-deployment cooldowns; content-policy-specific and context-window-specific fallback lists, separate from generic error fallbacks
PortkeyOpen-source gateway core; managed cloud plans on top250+ modelsPublishes its own overhead: "a total latency addition between 20-40ms compared to direct API calls." Reports serving over 25 million requests daily at 99.99% uptime, and is ISO 27001 and SOC 2 certified
Kong AI GatewayPlugin on Kong Gateway (open-source and enterprise tiers)OpenAI, Azure OpenAI, Bedrock, Anthropic, Gemini, Vertex AI, Cohere, Mistral, Hugging Face, Llama, xAI, DashScope, Cerebras, DeepSeek, Ollama, Databricks, vLLMRoute types map directly to OpenAI-shaped operations (llm/v1/chat, llm/v1/completions, llm/v1/embeddings, plus files, batches, assistants, and responses); records usage into Kong's existing logging plugins, while token-based limits come from a separate AI Rate Limiting Advanced plugin (AI Gateway Enterprise only)

The pattern across all three: none of them invented a new wire protocol. They normalize toward the OpenAI Chat Completions shape (or, for Kong, optionally skip normalization per its own docs) and add the policy layer around it.

The Tradeoff: You Just Added a New Single Point of Failure

A gateway sits between every call and every provider, which means a gateway outage is now an outage for every provider at once, even the ones that are individually healthy. This is the same tradeoff any middleware layer makes, and it is why deployment guides such as LiteLLM's production checklist have you run the gateway as multiple replicas behind readiness checks rather than a single process, and why it is worth keeping a tested direct-to-provider path in your client for when the gateway itself is unreachable. Portkey's own latency number, 20 to 40 milliseconds, is the cost of the normal path; an unavailable gateway is a different and larger cost, so factor gateway uptime into the reliability budget you were trying to improve by adding fallback in the first place.

When Adding One Is Actually Worth It

Skip a dedicated gateway for a single application calling a single provider; a direct client and a retry decorator cover that case with less operational surface. Reach for one once any of these is true: you call two or more providers and want one fallback path instead of hand-rolled try/except chains scattered across the codebase; you need spend visible by team or project rather than one combined bill; or several services or teams share the same provider credentials and you want one place to rotate a key, cap a budget, or cut off access, instead of redeploying every caller. All three are architectural reasons, not a specific traffic number, which is why "you'll know you need a gateway when configuration and cost tracking are duplicated in three different services" is a more honest threshold than any fixed request-per-day figure.

    Share Article:

  1. LI Test

  2. LI Test

Related resources.

How to Use the You.com Web Search API in TypeScript

How to Use the You.com Web Search API in TypeScript

September 22, 2026

Blog

How to Build a News Search Pipeline With the You.com Web Search API

How to Build a News Search Pipeline With the You.com Web Search API

September 22, 2026

Blog

What Is the You.com Web Search API? A Practical Guide for Developers

What Is the You.com Web Search API? Endpoint, Pricing, and Limits

September 22, 2026

Blog

How Authentication Works in the You.com Web Search API

How Authentication Works in the You.com Web Search API

September 21, 2026

Blog

How to Use Date Filters With the You.com Web Search API

How to Use Date Filters With the You.com Web Search API

September 21, 2026

Blog