Skip to content
Agent Guides10 min read

Running Claude Code with Other Models Through an Anthropic-Compatible Gateway

Tokens Team
Engineering
3 Oct 2026
On this page

Claude Code talks to exactly one kind of API: Anthropic Messages. Give it a different server that speaks the same format and it will drive whatever model sits behind that server. This page covers how to run Claude Code with other models through an Anthropic-compatible gateway, what stops working, and the handful of settings that decide whether the result is pleasant or painful.

bash
export ANTHROPIC_BASE_URL=https://tokens.bd
export ANTHROPIC_AUTH_TOKEN=$TOKENS_API_KEY
export ANTHROPIC_MODEL=deepseek/deepseek-v4.1-flash
claude

That is the whole trick. The rest of this post is about the edges.

How Claude Code picks its server and model#

Three environment variables do the work.

ANTHROPIC_BASE_URL replaces https://api.anthropic.com. Claude Code appends /v1/messages itself, so the value is the bare host. For Tokens that is https://tokens.bd, not https://tokens.bd/v1. Getting this wrong produces /v1/v1/messages and a 404 with code unsupported_endpoint, which is the most common first mistake.

ANTHROPIC_AUTH_TOKEN sends your key as Authorization: Bearer <key>. The alternative, ANTHROPIC_API_KEY, sends it as x-api-key and makes Claude Code ask for a one-time approval in interactive mode. Tokens reads either header, so use whichever you like; the auth token variant is the one with fewer prompts.

ANTHROPIC_MODEL sets the model for the session. On Tokens, model IDs are provider/model aliases. Check the exact ID on the model catalog or with GET /v1/models rather than guessing.

Claude Code also has a set of alias variables, and they matter more with a gateway than with Anthropic's own API:

VariableWhat it controls
ANTHROPIC_DEFAULT_OPUS_MODELThe model behind the opus alias
ANTHROPIC_DEFAULT_SONNET_MODELThe model behind the sonnet alias
ANTHROPIC_DEFAULT_HAIKU_MODELThe haiku alias, which Claude Code also uses for background tasks
CLAUDE_CODE_SUBAGENT_MODELThe model subagents run on
ANTHROPIC_CUSTOM_MODEL_OPTIONAdds one custom entry to the /model picker

If you leave the alias variables unset, a request for sonnet or haiku goes out with an Anthropic model name the gateway may not serve. And with ANTHROPIC_AUTH_TOKEN, background tasks run on your main model unless ANTHROPIC_DEFAULT_HAIKU_MODEL is set. That's a quiet cost leak: title generation and summaries billed at the price of whatever you picked for real work.

Put the settings in ~/.claude/settings.json#

Shell exports work, but the env block in ~/.claude/settings.json wins over them and survives new terminals. On Windows the file is %USERPROFILE%\.claude\settings.json.

/.claude/settings.json
{
  "env": {
    "ANTHROPIC_BASE_URL": "https://tokens.bd",
    "ANTHROPIC_AUTH_TOKEN": "tok_live_your_key",
    "ANTHROPIC_MODEL": "deepseek/deepseek-v4.1-flash",
    "ANTHROPIC_DEFAULT_OPUS_MODEL": "deepseek/deepseek-v4.1-flash",
    "ANTHROPIC_DEFAULT_SONNET_MODEL": "deepseek/deepseek-v4.1-flash",
    "ANTHROPIC_DEFAULT_HAIKU_MODEL": "deepseek/deepseek-v4.1-flash",
    "CLAUDE_CODE_SUBAGENT_MODEL": "deepseek/deepseek-v4.1-flash",
    "CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS": "1",
    "CLAUDE_CODE_MAX_CONTEXT_TOKENS": "128000"
  }
}

Keep the key out of your repo

This file is in your home directory, which is fine. A project-level .claude/settings.json is usually committed, so never put the key there. If you sync dotfiles to a public repo, exclude ~/.claude/settings.json or use the shell export instead.

The single-model version above is the safe starting point. Once it works, split the aliases: point opus at a stronger model for planning and hard refactors, and point haiku and the subagent model at something cheap. Most of the tokens Claude Code spends in a long session are re-sent context, not clever output, so moving the high-volume paths to a cheap model saves more than you'd expect.

The caveats Anthropic is clear about#

Anthropic doesn't support this setup#

Anthropic's gateway documentation says plainly that it "doesn't support routing Claude Code to non-Claude models through any gateway". This works, but nobody at Anthropic is testing your combination. A Claude Code update can change request shapes, and the gateway or the model has to cope. If something breaks right after an update, that's the first thing to suspect.

Claude Code sends Claude-specific fields to unknown models#

When Claude Code sees a model ID it doesn't recognise, it assumes the model supports every current Claude feature. It sends adaptive thinking, an effort setting (output_config), context_management, and beta tool fields.

CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS=1 strips most of these. It does not strip adaptive thinking or effort. Claude Code has a fallback: when the server rejects those fields with the error text it expects, it retries without them.

On Tokens, don't count on that fallback. When an upstream provider rejects a request with a 4xx, the gateway returns a generic error in Anthropic's shape (invalid_request_error) rather than passing the provider's original text through, so Claude Code may not recognise it as the "drop thinking and retry" case. The practical test: if a plain curl to the same model works (see the check at the end of this post) but every Claude Code turn fails with a 400, set CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS=1 if you haven't, and if it still fails, try a different model. Some routes tolerate the extra fields and some don't.

The /model picker won't list gateway models#

CLAUDE_CODE_ENABLE_GATEWAY_MODEL_DISCOVERY=1 makes Claude Code call GET /v1/models, but it keeps only IDs that contain claude or anthropic. A catalog of provider/model aliases like deepseek/deepseek-v4.1-flash gets filtered to nothing. Use ANTHROPIC_MODEL for the default and ANTHROPIC_CUSTOM_MODEL_OPTION to add one extra entry to the picker.

Context window and auto-compaction#

Claude Code assumes a 200K context window for any model it doesn't know. If your model's real window is smaller, set CLAUDE_CODE_MAX_CONTEXT_TOKENS to it. If it's larger, setting the real value lets you use it.

This matters more on Tokens than it might elsewhere, for the same reason as above. One way Claude Code learns it has run out of room is the upstream "prompt too long" error, and that error reaches you in generic form, so compaction triggered by the error may not fire. Set the window explicitly and Claude Code compacts on its own schedule instead. CLAUDE_CODE_AUTO_COMPACT_WINDOW (minimum 100,000) and CLAUDE_CODE_MAX_OUTPUT_TOKENS give you finer control.

Features that switch off#

Remote Control and voice dictation are disabled while a gateway credential is set. Fast mode checks api.anthropic.com directly, so it doesn't apply either. The tools themselves (reading files, editing, running commands) execute on your machine, so they don't depend on which model server you use.

Which models hold up inside Claude Code#

Claude Code is a demanding client. Nearly every step is a tool call: read a file, run a command, apply an edit. The system prompt and tool definitions are long, and every turn re-sends the whole conversation. That suggests a short list of properties to look for, roughly in order of importance.

Reliable tool calling with streamed arguments. A model that sometimes answers in prose when it should call a tool, or emits malformed JSON arguments, will stall the loop. This matters more than raw benchmark scores.

Exact-match edits. Claude Code edits files by exact string replacement. Models that paraphrase the "old" text while quoting it produce failed edits and retries. You'll see it quickly in practice; if a model keeps failing edits, switch.

Long context. 200K is comfortable. Many current models list 1M windows, which buys you longer sessions before compaction, though at the cost of re-sending more tokens per turn.

Sensible thinking behaviour. Some models always reason before answering. That's useful for planning and wasteful for "read this file". Models with a non-thinking mode, or with an effort setting, are easier to keep cheap.

Some current candidates, with list prices from their makers (checked October 2026):

ModelContextList price per 1M tokens (in / out)Notes
DeepSeek V4.1 Flash1M$0.30 / $1.20 peak, half that off-peakTools, thinking and non-thinking modes. DeepSeek's own API supports the Anthropic format.
GLM-5.31M$1.40 / $4.40Z.ai's coding and agent flagship. Thinking is always on.
Kimi K2.7 Code262,144$0.95 / $4.00Dedicated coding model, thinking only. tool_choice supports auto and none.
Qwen 3.8 Max1M$2.00 / $6.00Alibaba's strongest Qwen. Tools, image input.

A maker that ships its own Anthropic-format endpoint, as DeepSeek and Moonshot do, is a good sign for Claude Code compatibility. It isn't a guarantee, because the route Tokens uses for a model can differ from the maker's own API. What you'll pay through Tokens is on the model catalog and pricing, not in the table above. For more on picking a budget model, see cheap coding models that hold up, and for the long-context trade-off, million-token context for coding agents.

Set it up in one command with the Tokens CLI#

If you'd rather not edit JSON by hand, the Tokens CLI writes the Claude Code settings for you. It's a single dependency-free Node script (Node 18+); read it before you run it.

bash
curl -fsSL https://tokens.bd/cli/tokens.mjs -o tokens.mjs
node tokens.mjs setup --base-url https://tokens.bd --agents claude

On Windows PowerShell:

powershell
iwr https://tokens.bd/cli/tokens.mjs -OutFile tokens.mjs; node tokens.mjs setup --base-url https://tokens.bd --agents claude

Without --key, it opens your browser so you can approve a short code at /dashboard/connect/cli, which creates a new key named CLI (<label>). It lists the models your account can use, asks you to pick a default, then merges three values into ~/.claude/settings.json: ANTHROPIC_BASE_URL, ANTHROPIC_AUTH_TOKEN and ANTHROPIC_MODEL. The existing file is copied to settings.json.tokens-backup-<timestamp> first. If the file has comments or anything else it can't parse safely, the CLI leaves it alone and points you to the copy-and-paste snippet instead.

The CLI writes only those three variables. Add the alias variables, the subagent model and CLAUDE_CODE_MAX_CONTEXT_TOKENS yourself afterwards. The full walkthrough, including --dry-run, is in the Tokens CLI docs.

To go back to your Anthropic subscription, delete the env entries or copy the backup file over settings.json.

Check that Claude Code is really using the gateway#

Test the endpoint directly first. This is the request from Claude Code's gateway docs with Tokens values; it asks for one output token, so it costs next to nothing:

bash
curl -sS -w '\n%{http_code}\n' -X POST "https://tokens.bd/v1/messages" \
  -H "Authorization: Bearer $TOKENS_API_KEY" \
  -H "anthropic-version: 2023-06-01" -H "content-type: application/json" \
  -d '{"model": "deepseek/deepseek-v4.1-flash", "max_tokens": 1, "messages": [{"role": "user", "content": "."}]}'

A 200 means the key, the model and the route all work. Then start claude and run /status. The "Anthropic base URL" line should show https://tokens.bd, and the auth line should name ANTHROPIC_AUTH_TOKEN. claude doctor checks the settings files for mistakes.

If the curl fails, the status code tells you where: 401 is the key, 403 model_not_allowed_on_key or tier_permission_denied is the model choice, 402 is your balance. Reading gateway errors goes through each one.

If you don't have a key yet, create one at API keys with a monthly spend cap, since a long agent session is exactly the kind of workload a cap is for. The Claude Code docs page has the same settings in reference form.

Third-party details (Claude Code variables and behaviour, model specs and list prices) checked on 2026-10-03.

Sources: https://code.claude.com/docs/en/llm-gateway-connect · https://code.claude.com/docs/en/llm-gateway-protocol · https://code.claude.com/docs/en/model-config · https://api-docs.deepseek.com/quick_start/pricing · https://docs.z.ai/guides/overview/pricing · https://platform.kimi.ai/docs/pricing/chat · https://www.alibabacloud.com/help/en/model-studio/model-pricing

Was this page helpful?

Still stuck? Open a support ticket

Use the coding models you already know, through one API

One key for OpenAI- and Anthropic-compatible tools. Pay in BDT or USD, and keep the coding agent you already use.

Create an account