# Claude Agent SDK

> Run agents built with the Claude Agent SDK (Python and TypeScript) on Tokens: set ANTHROPIC_BASE_URL and a Tokens key, choose a model id, handle the opus, sonnet and haiku aliases, stream and add tools.

The Claude Agent SDK is Anthropic's library for running the Claude Code agent loop inside your own Python or TypeScript program: built-in tools to read and edit files and run commands, permissions, sessions, subagents and MCP. It is not a thin API client. The SDK starts a Claude Code process and talks to it, and that process sends Anthropic Messages requests to whatever `ANTHROPIC_BASE_URL` says. Pointed at `https://tokens.bd`, those requests go to Tokens' [/v1/messages](/docs/messages) endpoint. If you only need to call a model, the [Anthropic SDK](/docs/anthropic-sdk) is the smaller choice.

:::note[Checked against the documentation]
Based on the Claude Agent SDK documentation at code.claude.com, checked October 2026 against `claude-agent-sdk` 0.2.165 for Python and `@anthropic-ai/claude-agent-sdk` 0.3.296 for TypeScript. The code was checked against the documentation, not run end to end against Tokens.
:::

:::warning[Non-Claude models are best effort]
Anthropic's gateway documentation says it "doesn't support routing Claude Code to non-Claude models through any gateway". The Agent SDK runs the same agent as Claude Code, so the same applies. It works when the model handles tool calls well, and that is per model. If an agent loses its way in long runs, switch models before debugging the SDK. [Claude Code](/docs/claude-code) has the same caveat.
:::

## What you need

- A Tokens key from [API keys](/docs/api-keys), exported as `TOKENS_API_KEY`.
- A model id from [/models](/models), for example `deepseek/deepseek-v4.1-flash`, that supports tool calling ([Choosing a model](/docs/choosing-a-model)).
- Python 3.10 or newer, or Node.js 18 or newer. Both packages bundle a native Claude Code binary, so most installs need no separate Claude Code install. Some do not: a pip source install (for example on Windows on ARM), or an npm install that skips optional dependencies (`npm ci --omit=optional`). In those cases install Claude Code natively ([Claude Code](/docs/claude-code)), and in TypeScript set `pathToClaudeCodeExecutable`.

```bash
export TOKENS_API_KEY="tok_live_your_key"
```

## How the SDK reaches Tokens

The SDK has no gateway options of its own. Anthropic's documentation says it passes environment variables to the Claude Code process it starts, and each SDK has an `env` option for that. You set three things:

| Setting                                          | Value                                                                                                 |
| ------------------------------------------------ | ----------------------------------------------------------------------------------------------------- |
| `ANTHROPIC_BASE_URL`                             | `https://tokens.bd`, without `/v1`. The process adds `/v1/messages` itself.                      |
| `ANTHROPIC_AUTH_TOKEN`                           | Your Tokens key. Sent as `Authorization: Bearer`, which Tokens accepts.                               |
| The model (`model` option and the alias variables) | A Tokens model id such as `deepseek/deepseek-v4.1-flash`, not `sonnet` or `opus`. See the next section.        |

`ANTHROPIC_API_KEY` also works (it is sent as `x-api-key`). Set only one of the two. This page uses `ANTHROPIC_AUTH_TOKEN`, the same as [Claude Code](/docs/claude-code).

The two SDKs treat `env` differently, and the difference matters:

- **TypeScript.** The process inherits your environment by default, but setting `options.env` replaces it entirely. Spread `process.env` into it, or the process loses `PATH` and everything else.
- **Python.** `ClaudeAgentOptions(env=...)` is merged on top of the inherited environment.

The SDK does not read `.env` files. Load them yourself before you start the SDK, for example with `dotenv`.

## Choose the model id and fix the aliases

The `model` option takes "a Claude model alias or full model name" (`ClaudeAgentOptions.model` in Python, `options.model` in TypeScript). Pass a Tokens model id and the process sends it to Tokens as written.

Aliases are the catch. Anthropic's model documentation says `sonnet`, `opus` and `haiku` resolve to the latest Claude model for your provider, and that `ANTHROPIC_BASE_URL` "changes where requests are sent, not which model answers them". So with Tokens, an alias still resolves to a Claude model name unless you redirect it, and that is not necessarily a model Tokens serves. Parts of the agent use aliases without you asking:

| What                                                                              | Which model it uses                                                  |
| --------------------------------------------------------------------------------- | -------------------------------------------------------------------- |
| Your main loop                                                                    | The `model` option                                                   |
| Background work such as session titles                                            | The `haiku` alias, or `ANTHROPIC_DEFAULT_HAIKU_MODEL` when it is set |
| `sonnet`, `opus` and `haiku` when you or a subagent definition name them          | `ANTHROPIC_DEFAULT_SONNET_MODEL`, `ANTHROPIC_DEFAULT_OPUS_MODEL`, `ANTHROPIC_DEFAULT_HAIKU_MODEL` |
| Subagents with no model of their own                                              | `CLAUDE_CODE_SUBAGENT_MODEL`                                         |

Set all of them to the same Tokens id, as the [Claude Code page](/docs/claude-code) does, so no part of the agent asks for a model Tokens does not have. Each variable has to be a full model id. Copy ids from [/models](/models) or `GET https://tokens.bd/v1/models`.

Two more variables from the Claude Code page also apply, because the SDK runs the same process:

- `CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS=1` removes most Claude-only beta fields from requests, which prevents `400` errors from non-Claude models.
- `CLAUDE_CODE_MAX_CONTEXT_TOKENS` tells Claude Code the real context window. For ids it does not recognize it assumes 200K tokens, so a model with a smaller window fails with "prompt too long" instead of compacting. Take the number from the model's page in [/models](/models).

## Python: a minimal agent

```bash
pip install --upgrade claude-agent-sdk
```

```python title="agent.py"
import asyncio
import os

from claude_agent_sdk import AssistantMessage, ClaudeAgentOptions, ResultMessage, query

MODEL = "deepseek/deepseek-v4.1-flash"

ENV = {
    "ANTHROPIC_BASE_URL": "https://tokens.bd",
    "ANTHROPIC_AUTH_TOKEN": os.environ["TOKENS_API_KEY"],
    "ANTHROPIC_DEFAULT_OPUS_MODEL": MODEL,
    "ANTHROPIC_DEFAULT_SONNET_MODEL": MODEL,
    "ANTHROPIC_DEFAULT_HAIKU_MODEL": MODEL,
    "CLAUDE_CODE_SUBAGENT_MODEL": MODEL,
    "CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS": "1",
}

options = ClaudeAgentOptions(
    model=MODEL,
    allowed_tools=["Read", "Glob", "Grep"],  # read-only tools, approved automatically
    max_turns=8,
    setting_sources=[],  # ignore ~/.claude/settings.json, see the note below
    env=ENV,
)


async def main() -> None:
    async for message in query(
        prompt="List the Python files in this folder and say in one sentence what each one does.",
        options=options,
    ):
        if isinstance(message, AssistantMessage):
            for block in message.content:
                if hasattr(block, "text"):
                    print(block.text)
                elif hasattr(block, "name"):
                    print(f"[tool: {block.name}]")
        elif isinstance(message, ResultMessage):
            print(f"Done: {message.subtype}, {message.num_turns} turns")


asyncio.run(main())
```

`query()` returns an async iterator. Each item is a message: the model's text, a tool call, a tool result, then a final `ResultMessage`. `allowed_tools` only approves tools without a prompt. It does not remove the others. To remove a tool, use `disallowed_tools`.

:::note[Settings files can override your variables]
Claude Code reads `~/.claude/settings.json` and project settings. Anthropic's documentation says that when a shell export and a settings-file `env` block set the same variable, the settings-file value wins. If you already set up [Claude Code](/docs/claude-code) on the same machine, that file's base URL, key or model would apply to your agent. `setting_sources=[]` (Python) skips those files. It also means the agent does not load `CLAUDE.md`, skills or project settings. Leave the line out if you want them.
:::

## TypeScript: a minimal agent

```bash
npm install @anthropic-ai/claude-agent-sdk
npm install --save-dev tsx
```

Set `"type": "module"` in `package.json` so top-level `await` works, or name the file `agent.mts`.

```ts title="agent.ts"
import { query } from "@anthropic-ai/claude-agent-sdk";

const model = "deepseek/deepseek-v4.1-flash";

const env = {
  ...process.env, // required: setting env replaces the whole environment
  ANTHROPIC_BASE_URL: "https://tokens.bd",
  ANTHROPIC_AUTH_TOKEN: process.env.TOKENS_API_KEY,
  ANTHROPIC_DEFAULT_OPUS_MODEL: model,
  ANTHROPIC_DEFAULT_SONNET_MODEL: model,
  ANTHROPIC_DEFAULT_HAIKU_MODEL: model,
  CLAUDE_CODE_SUBAGENT_MODEL: model,
  CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS: "1",
};

for await (const message of query({
  prompt: "List the TypeScript files in this folder and say in one sentence what each one does.",
  options: {
    model,
    allowedTools: ["Read", "Glob", "Grep"], // read-only tools, approved automatically
    maxTurns: 8,
    settingSources: [], // ignore ~/.claude/settings.json
    env,
  },
})) {
  if (message.type === "assistant" && message.message?.content) {
    for (const block of message.message.content) {
      if ("text" in block) console.log(block.text);
      else if ("name" in block) console.log(`[tool: ${block.name}]`);
    }
  } else if (message.type === "result") {
    console.log(`Done: ${message.subtype}`);
  }
}
```

```bash
npx tsx agent.ts
```

If you would rather not pass `env` in code, export the same variables in the shell that runs your program. In TypeScript the process inherits them. In Python they pass through as well, since `env` is merged on top of the environment.

## Stream the output

By default the SDK yields a whole text block or tool call after the model finishes it. For token-by-token text, turn on partial messages. The SDK then also yields raw Messages API stream events, and you read the text deltas.

:::code-tabs

```python title="Python"
from claude_agent_sdk import ClaudeAgentOptions, query
from claude_agent_sdk.types import StreamEvent

# MODEL and ENV are defined in the minimal program above
options = ClaudeAgentOptions(
    model=MODEL,
    include_partial_messages=True,
    setting_sources=[],
    env=ENV,
)


async def stream() -> None:
    async for message in query(
        prompt="Explain optimistic locking in two sentences.", options=options
    ):
        if isinstance(message, StreamEvent):
            event = message.event
            if event.get("type") == "content_block_delta":
                delta = event.get("delta", {})
                if delta.get("type") == "text_delta":
                    print(delta.get("text", ""), end="", flush=True)
```

```ts title="TypeScript"
// model and env are defined in the minimal program above
for await (const message of query({
  prompt: "Explain optimistic locking in two sentences.",
  options: { model, includePartialMessages: true, settingSources: [], env },
})) {
  if (message.type === "stream_event") {
    const event = message.event;
    if (event.type === "content_block_delta" && event.delta.type === "text_delta") {
      process.stdout.write(event.delta.text);
    }
  }
}
```

:::

The events have the Anthropic streaming shape, which Tokens produces for any model ([Messages](/docs/messages), [Streaming](/docs/streaming)). Stream events cover the main agent only. Token deltas from subagents are not forwarded.

## Add your own tools

Besides the built-in tools, you can give the agent functions from your own code as an in-process MCP server. In Python, `@tool` and `create_sdk_mcp_server` build it. In TypeScript, `tool` and `createSdkMcpServer` do, with a Zod schema.

:::code-tabs

```python title="Python"
from typing import Any

from claude_agent_sdk import ClaudeAgentOptions, create_sdk_mcp_server, tool


@tool("get_weather", "Current weather for a city", {"city": str})
async def get_weather(args: dict[str, Any]) -> dict[str, Any]:
    return {"content": [{"type": "text", "text": f"{args['city']}: light rain, 29C"}]}


weather = create_sdk_mcp_server(name="weather", version="1.0.0", tools=[get_weather])

# MODEL and ENV are defined in the minimal program above
options = ClaudeAgentOptions(
    model=MODEL,
    mcp_servers={"weather": weather},
    allowed_tools=["mcp__weather__get_weather"],  # MCP tools are named mcp__<server>__<tool>
    max_turns=5,
    setting_sources=[],
    env=ENV,
)
```

```ts title="TypeScript"
import { createSdkMcpServer, query, tool } from "@anthropic-ai/claude-agent-sdk";
import { z } from "zod";

const getWeather = tool(
  "get_weather",
  "Current weather for a city",
  { city: z.string() },
  async ({ city }) => ({
    content: [{ type: "text", text: `${city}: light rain, 29C` }],
  }),
);

const weather = createSdkMcpServer({ name: "weather", version: "1.0.0", tools: [getWeather] });

// model and env are defined in the minimal program above
const options = {
  model,
  mcpServers: { weather },
  allowedTools: ["mcp__weather__get_weather"],
  maxTurns: 5,
  settingSources: [],
  env,
};
```

:::

Pass `options` to `query()` as in the minimal examples. The TypeScript package needs `zod` 4 as a peer dependency. Each tool round is another `/v1/messages` request, so an agent run costs several requests and counts against your [rate limits](/docs/rate-limits). See [Tool calling](/docs/tool-calling) for how tool calls are handled for non-Claude models.

## Check that it works

First test the endpoint and the model id without the SDK (macOS, Linux, WSL or Git Bash):

```bash
curl -sS -w '\n%{http_code}\n' -X POST "https://tokens.bd/v1/messages" \
  -H "Authorization: Bearer $TOKENS_API_KEY" \
  -H "anthropic-version: 2023-06-01" -H "content-type: application/json" \
  -d '{"model": "deepseek/deepseek-v4.1-flash", "max_tokens": 1, "messages": [{"role": "user", "content": "."}]}'
```

A `200` means the key and model are fine. Then run the minimal program. You should see the tool names it calls, the model's answer and `Done: success`. The requests also show up in Usage analytics in the [dashboard](/dashboard), under the model id you set. If the log lists a different model, one of the alias variables was not set.

`ResultMessage` has a `total_cost_usd` field. That is the SDK's own estimate, not what Tokens bills. Tokens' [usage page and `GET /v1/tokens/usage`](/docs/models-and-usage) are the source of truth.

## Choosing a model

The agent loop relies on tool calls, long contexts and following a system prompt. Models differ a lot here. [Choosing a model](/docs/choosing-a-model) covers which suit agent work. Use `fallback_model` (Python) or `fallbackModel` (TypeScript) only with another Tokens id, never a Claude name. Lowering `max_turns` and setting a monthly spend cap on the key ([API keys](/docs/api-keys)) limits the damage from a runaway loop.

## Limits and what does not work

- **Claude-only features.** Fast mode checks Anthropic's API directly and does not work through Tokens. Remote Control and voice dictation are disabled while a gateway credential is set. Extended thinking and prompt caching work only where the model behind the Tokens id supports them ([Reasoning](/docs/reasoning), [Prompt caching](/docs/prompt-caching)).
- **Claude Code's own model picker.** The `/model` picker belongs to the interactive CLI. In the SDK you set the model in options.
- **Server-side Anthropic tools.** Tools that run on Anthropic's side, such as hosted web search, are not documented for gateways, and Anthropic does not document them for non-Claude models. Test one before you rely on it, and prefer your own tools or MCP servers.
- **Login.** Anthropic does not allow third-party developers to offer claude.ai login or its rate limits in products built on the SDK, so use an API key as shown here.
- **Cloud provider modes.** Variables such as `CLAUDE_CODE_USE_BEDROCK` or `CLAUDE_CODE_USE_VERTEX` select other providers. Do not set them with Tokens.
- **Background traffic.** Claude Code sends version checks and telemetry to Anthropic outside the gateway. `CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1` turns that off, at the cost of auto-updates.

## Troubleshooting

**404 on every request.** The base URL ends in `/v1`. The process appends `/v1/messages` itself, so use `https://tokens.bd`.

**401 `invalid_api_key`, or `Not logged in`.** The key is wrong, or the variables never reached the process. In TypeScript, check that `env` has `...process.env`. In Python, check the `env` dict. Also check that a `~/.claude/settings.json` is not overriding your values (use `settingSources: []`). The SDK does not load `.env` files.

**404 `model_not_found`.** The id is not one Tokens knows, or an alias (`sonnet`, `opus`, `haiku`) resolved to a Claude name. Set the alias variables to your Tokens id, and check the id against `GET https://tokens.bd/v1/models`.

**400 errors about `thinking`, `effort`, `context_management` or unknown fields.** Claude Code sends adaptive thinking and beta fields that non-Claude models reject. `CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS=1` removes most of them. If one model still fails, try another.

**"Prompt is too long" or the session never compacts.** Claude Code assumes a 200K context for unknown ids. Set `CLAUDE_CODE_MAX_CONTEXT_TOKENS` to the model's real window.

**The process fails to start.** No bundled binary was installed. Install Claude Code natively, and in TypeScript set `pathToClaudeCodeExecutable`.

**402 `insufficient_credits`, 403 `model_not_allowed_on_key` or `tier_permission_denied`, 429 `window_exhausted` or `concurrency_limit`.** These are account limits, not SDK problems. A run with several subagents sends requests in parallel, which can hit `concurrency_limit`. See [Errors](/docs/errors) and [Troubleshooting](/docs/troubleshooting), and include the `x-tokens-request-id` header value when you contact [support](/docs/support).

---
Page: https://tokens.bd/docs/claude-agent-sdk
