TL;DR: An AI gateway is a single OpenAI-compatible API that sits in front of many model providers. It adds model routing, automatic fallbacks, spend caps, caching, and one bill. The leading hosted gateways charge no inference markup — they pass provider list prices through and earn on credit-purchase fees or plans instead. You need one when you run more than one model, share API keys across a team, or can't afford to go down with your provider. You don't when it's one model, one key, and low volume.
Your AI feature ships. It calls one provider, with one key, and it works — until 14:00 UTC on 29 September 2026, when Anthropic's API started returning errors for everyone, for 59 minutes straight. On 25 September, OpenAI had its own 56-minute outage. On 3 September, ChatGPT, Claude, and Grok all went down within a 90-minute window when a regional Azure failure took out East US. Your uptime is your provider's uptime — unless something sits between your app and the provider that can route around the failure.
That something is an AI gateway. This post explains what it is, what it does for you, what it costs, what the landscape looks like in October 2026, and — honestly — when you can skip it.
What an AI gateway actually is#
Strip away the marketing and an AI gateway is a routing and control layer between your application and model providers. It normalizes the different provider APIs behind one interface — almost always OpenAI-compatible — and adds the operational layer production AI apps need: unified key management, model routing, automatic retries and fallbacks, caching, rate limits, budgets and spend caps, guardrails, and observability (logs, traces, cost per key or team).
The mental model that sticks: it's what an API gateway (Kong, Nginx, Cloudflare) has always been for HTTP services, applied to model providers. One endpoint in, many providers out, with policy, resilience, and billing handled at the layer in between.
This is also why OpenAI SDK compatibility became the standard interface. Change one base_url and your code talks to dozens of models instead of one — no SDK swaps, no request-format rewrites. We walk through that one-line change in Use Any LLM With the OpenAI SDK.
The five jobs a gateway does for you#
1. One API for every model#
Each provider has its own SDK, auth scheme, and request format. A gateway collapses that into one OpenAI-format API and one API key. You stop writing provider adapters and start swapping models with a string change. OpenRouter lists 500+ models from 80+ providers behind its endpoint; Portkey advertises 1,600+ models across 40+ providers; tokens.bd offers 62 models from 10 providers behind one key. The value isn't just convenience — it's optionality. When a better or cheaper model launches, switching is a config change, not a migration.
2. Automatic fallbacks when a provider goes down#
This is the job that pays for the whole layer at 3 AM. Gateways retry with backoff and route to the next provider on 429s, 5xx errors, and connection failures. The 2026 outage record makes the case better than any benchmark:
- 29 September 2026: Anthropic API returned errors for everyone, 14:00–14:59 UTC (59 minutes).
- 25 September 2026: OpenAI outage, 22:59–23:55 UTC (56 minutes).
- 3 September 2026: ChatGPT, Claude, and Grok went down within about 90 minutes of each other — a regional failure in Microsoft Azure's East US region. Gemini, hosted on Google Cloud, was largely unaffected.
- 16 August 2026: claude.ai, the Anthropic API, Claude Code, and Cowork all failed together — switching surfaces was no fallback.
Independent monitoring put the major providers at 99.9%+ availability in 2026 — which still allows roughly nine hours of downtime per year. With a gateway and a fallback chain (e.g., Claude → GPT → Gemini), a single-provider outage becomes a latency blip instead of an incident.
One real-world gotcha worth knowing: LiteLLM's default retry config once spent about 4.7 seconds per request retrying a dead primary before falling back; setting num_retries: 0 brought fallback time to 46–87ms. Fallbacks only work if you configure them aggressively.
3. Cost control: one bill, caching, spend caps#
Direct integrations scatter spend across N provider bills with N dashboards. A gateway gives you one bill and the controls that keep it small:
- Spend caps and budgets per key, per team, or per customer. tokens.bd lets you set a monthly cap when you create a key — a runaway agent loop stops at the cap — with email alerts at 50%, 75%, and 90%.
- Caching of repeated prompts, so identical or semantically similar requests don't bill twice (see prompt caching explained).
- Rate limits that protect you from your own bugs before they become invoices.
4. Key governance and privacy#
Handing every developer and every service a raw provider key is how secrets leak into repos and how one compromised laptop becomes an unlimited bill. Gateways issue virtual keys — scoped, revocable, capped — while the real provider keys live in the gateway's vault. Privacy-wise, the better gateways are explicit: tokens.bd states it keeps only what billing needs (model, tokens, cost, latency) and never stores prompts or responses, with no training on your data.
5. Observability#
Which model served which request, how much each key spent, where the latency went — gateways log all of it in one place. When something returns a weird error, reading gateway errors is the debugging skill that replaces guessing across five provider dashboards.

Do gateways slow you down or mark up prices?#
Two honest questions, two verified answers.
Latency. A well-deployed proxy adds real but small overhead — the rule of thumb from a 2026 architecture decision doc is under 10ms versus direct calls. LiteLLM's own docs report about 2ms median overhead on a scaled-out setup. The exception is policy-heavy paths: Cloudflare AI Gateway's guardrails cost around 500ms and buffer streamed responses. For chat and coding agents, that's noise; for latency-critical single-request paths, it's a reason to measure.
Pricing. The leading hosted gateways charge zero inference markup — provider list prices pass through untouched. OpenRouter, Vercel AI Gateway, and Opper all do this. They monetize differently: OpenRouter charges 5.5% on credit purchases on the Standard plan and 8% on Business (the tier with EU/US in-region routing); Cloudflare's unified billing takes 5% on credits; Portkey and Kong sell platform plans. Self-hosted gateways (LiteLLM, Kong, Bifrost) are free software — you pay with your own infrastructure and operations time.
The upshot: for most developers, a gateway doesn't make tokens more expensive. It makes spending controllable.
The 2026 landscape: five gateways to know#
| Gateway | Type | Scale (as claimed, Oct 2026) | Pricing model |
|---|---|---|---|
| OpenRouter | Hosted marketplace | 500+ models, 80+ providers | No inference markup; 5.5% fee on credit buys (Standard), 8% (Business) |
| Portkey | Hosted + open-source core (MIT) | 1,600+ models, 40+ providers | Free Developer (10k logs/mo); Production from $49/mo |
| LiteLLM | Open-source Python (MIT), self-hosted | 100+ provider integrations; ~59k GitHub stars | Free software; enterprise quote-based |
| Cloudflare AI Gateway | Managed on Cloudflare's network | Caching, fallbacks, guardrails, DLP | Core free; 5% fee on unified-billing credits |
| Kong AI Gateway | AI plugins on Kong (OSS + Enterprise) | Multi-provider routing, MCP + agent-to-agent | OSS free; managed from $100/mo per LLM model |
A few 2026 developments worth knowing:
- Kong AI Gateway 2.0 went GA on 1 September 2026 as its own product — separate runtime, control plane, and release cycle — covering LLM, MCP, and agent-to-agent traffic.
- Bloomberg reported on 16 August 2026 that Stripe agreed to acquire OpenRouter for over $7B (about 5x its $1.3B Series B valuation from May 2026). Neither company has confirmed it, and the deal was not closed as of October 2026 — treat it as reported, not done.
- Portkey raised a $15M Series A on 19 February 2026, led by Elevation Capital with Lightspeed, citing 500B+ tokens and 125M requests/day across 24,000+ organizations.
- Security note: on 24 March 2026, malicious LiteLLM versions 1.82.7 and 1.82.8 sat on PyPI for about 40 minutes carrying a credential stealer. The fix is evergreen: pin your dependency versions.
- Cloudflare merged AI Gateway with Workers AI into a single control plane in August 2026.
For a head-to-head with the most popular option, see tokens.bd vs OpenRouter — including what OpenRouter does better (breadth, free models, routing control) and where tokens.bd differs (BDT billing, per-key spend caps, one-dashboard support).

When you actually need one — and when you don't#
You need a gateway when:
- You use more than one model family — writing and maintaining N provider integrations is toil.
- Multiple people or services share LLM spend — you need budgets, rate limits, and audit trails per key.
- Downtime costs you — a provider outage shouldn't be your outage. Fallback chains are the cheapest uptime insurance in AI infrastructure.
- You want central prompt logging, cost analytics, or PII guardrails across everything.
- Paying the provider directly is hard for you — for example, developers in Bangladesh, where international-card online payments are capped at USD 300 per transaction (Bangladesh Bank FE Circular 26, 2019) with paperwork each time.
You can skip it when:
- It's one model, one key, low volume — a weekend prototype or a single cron job.
- You're on a latency-critical single-request path where even 10ms matters and you can tolerate the provider's uptime.
- You genuinely want per-request provider ordering, price sorting, or routing flags — power features some gateways have and simpler ones don't.
The common middle ground: start direct, and move to a gateway the week you add a second model or a second developer. That week arrives sooner than most teams expect.
Try one in five minutes#
The whole pitch of a gateway is that integration is boring. With tokens.bd, it's one environment variable — the same key works on the OpenAI-compatible endpoint and the Anthropic-compatible endpoint:
from openai import OpenAI
client = OpenAI(
base_url="https://tokens.bd/v1",
api_key="tk_your_key_here", # create one at tokens.bd, set a monthly cap
)
resp = client.chat.completions.create(
model="qwen/qwen-3.7-flash", # or "anthropic/claude-sonnet-5.5"
messages=[{"role": "user", "content": "Explain idempotency in one paragraph."}],
)
print(resp.choices[0].message.content)Swap the model string and you're on a different provider — Claude Sonnet 5.5 at $2.20/$11.00 per million tokens, Qwen 3.7 Flash at $0.03/$0.14, GPT-6 Luna at $0.11/$0.55. The full 62-model catalog shows every price in USD and BDT. If the call fails, the gateway's automatic failover moves the request to the next upstream before the error reaches you.
The Bangladesh angle#
Gateways solve a very local problem for Bangladeshi developers: paying for AI APIs at all. International cards are hard to get and capped per transaction; tokens.bd lets you pay in USD or BDT by invoice — no international card needed — with a printable receipt for every payment. Combine that with per-key monthly spend caps (a runaway agent stops at the cap, with alerts at 50/75/90%), BDT-denominated prices on every model in the catalog, and one dashboard for keys, receipts, and support tickets, and the "gateway" stops being an abstraction — it's the difference between shipping and not shipping.
AI gateway FAQ#
What is an AI gateway in simple terms?
It's a single API endpoint that routes your requests to many AI model providers. Instead of integrating OpenAI, Anthropic, Google, and DeepSeek separately, you point your code at the gateway (usually via one base_url change) and pick models with a string. The gateway handles key management, retries, fallbacks, caching, spend limits, and logging.
Is an AI gateway the same as an API gateway? Same idea, different traffic. A classic API gateway (Kong, Nginx) routes and protects HTTP services; an AI gateway does that for model providers, adding AI-specific jobs: model routing, provider failover, prompt caching, token-level cost tracking, and guardrails.
Do AI gateways charge extra per token? The leading hosted ones don't. OpenRouter, Vercel AI Gateway, and Cloudflare pass provider list prices through with no inference markup; they charge fees on credit purchases (OpenRouter 5.5–8%, Cloudflare 5%) or platform plans. Self-hosted gateways like LiteLLM are free software — your cost is infrastructure and operations.
Do gateways add noticeable latency? Usually not: a well-deployed proxy adds under 10ms. The exceptions are policy-heavy paths — e.g., Cloudflare's guardrails add around 500ms and buffer streaming — and misconfigured retries, which can delay failover by seconds. Measure on your own traffic.
When should I use an AI gateway? When you use more than one model, share keys across a team, need spend caps and audit trails, or can't afford a provider outage (2026 saw 56–59-minute outages at both OpenAI and Anthropic). Skip it for one-model, low-volume prototypes where direct API calls are simpler.
Facts and prices checked 9 October 2026. Gateway plans, fees, and model counts change; the provider's page is the authority, and the tokens.bd catalog is the authority for prices on Tokens.
Sources: OpenRouter pricing · OpenRouter business tier · Vercel AI Gateway vs OpenRouter · OpenRouter alternatives / Stripe report · Portkey raise · Top 5 enterprise AI gateways 2026 · Best LLM gateways 2026 comparison · Anthropic Sept 29 outage · Provider uptime comparison · Sept 3 multi-provider outage · Why your AI app needs a gateway layer · tokens.bd vs OpenRouter



