# OpenAI Agents SDK

> Run agents built with the OpenAI Agents SDK (Python and TypeScript) on Tokens: Chat Completions model class, the slash-in-model-id problem, tracing, tools and streaming.

The OpenAI Agents SDK is OpenAI's library for building agents: an `Agent` with instructions and tools, run by a `Runner`. It talks to models through the OpenAI API, so it can reach Tokens at `https://tokens.bd/v1`. Two defaults get in the way, and both are easy to fix: the SDK calls the Responses API unless you tell it otherwise, and it reads a model id that contains a slash as `provider/model`. Every Tokens model id contains a slash, so this page shows how to set both up.

:::note[Checked against the documentation]
Based on the OpenAI Agents SDK documentation, checked October 2026 against the Python package `openai-agents` 0.23.1 (released 2 October 2026) and the TypeScript package `@openai/agents` 0.20.0. The code was checked against the documentation and the SDK's own examples, not run end to end against Tokens. The SDK changes quickly, so compare with the [official docs](https://openai.github.io/openai-agents-python/models/) if something here no longer matches your version.
:::

## What you need

- A Tokens key from [API keys](/docs/api-keys), exported as `TOKENS_API_KEY`.
- A model id from [/models](/models), for example `deepseek/deepseek-v4.1-flash`. Agents call tools, so pick a model that supports tool calling ([Choosing a model](/docs/choosing-a-model)).
- Python 3.10 or newer for the Python package. The TypeScript package needs the `openai` package version 7.2 or newer when you pass your own client.

```bash
export TOKENS_API_KEY="tok_live_your_key"
```

## Why the defaults need changing

| Default in the SDK                                    | What happens with Tokens                                                                                                                             | Fix                                                                          |
| ----------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------- |
| Uses the Responses API                                | Tokens passes `/v1/responses` through, but it only works if the provider behind the model implements it ([Responses](/docs/responses)). Chat Completions works for every model. | Use `OpenAIChatCompletionsModel`, or `set_default_openai_api("chat_completions")`. |
| Reads `prefix/name` as a provider prefix              | The SDK documents that unknown prefixes raise `UserError`, and that `openai/...` is shortened to the part after the slash. A Tokens id such as `deepseek/deepseek-v4.1-flash` hits this rule. | Pass a model object, use a custom `ModelProvider`, or switch the prefix modes. |
| Uploads traces to OpenAI                              | Traces go to OpenAI's servers and need an OpenAI key. With a Tokens key the upload fails with a 401 in your logs.                                    | Turn tracing off.                                                            |

The SDK's documentation lists these under its "non-OpenAI models" section. It recommends Chat Completions when a provider lacks Responses support, and says unknown prefixes raise `UserError` instead of being passed through.

## Python: use a Chat Completions model object

This is the setup to start with. You create an `AsyncOpenAI` client that points at Tokens, wrap it in `OpenAIChatCompletionsModel`, and give that object to the agent. Because the agent receives a model object and not a model name string, the SDK does not parse the model id, so the slash is not an issue.

```bash
pip install --upgrade openai-agents
```

```python title="agent.py"
import asyncio
import os

from openai import AsyncOpenAI
from agents import Agent, OpenAIChatCompletionsModel, Runner, set_tracing_disabled

set_tracing_disabled(True)  # traces would go to OpenAI; see the Tracing section

client = AsyncOpenAI(
    base_url="https://tokens.bd/v1",
    api_key=os.environ["TOKENS_API_KEY"],
)

agent = Agent(
    name="Reviewer",
    instructions="You are a concise senior engineer. Answer in at most three sentences.",
    model=OpenAIChatCompletionsModel(model="deepseek/deepseek-v4.1-flash", openai_client=client),
)


async def main() -> None:
    result = await Runner.run(agent, "When should I use a dataclass instead of a dict?")
    print(result.final_output)


asyncio.run(main())
```

`base_url` keeps the `/v1` at the end. The model string goes to Tokens exactly as written, so it must match an id from [/models](/models) or `GET https://tokens.bd/v1/models`.

## Python: set a model for every agent in one place

If you have many agents, a custom `ModelProvider` passed through `RunConfig` applies one model setup to a whole run, and your `Agent` objects can keep plain model names. This follows the SDK's own `custom_example_provider.py` example.

```python title="provider.py"
import asyncio
import os

from openai import AsyncOpenAI
from agents import (
    Agent,
    Model,
    ModelProvider,
    OpenAIChatCompletionsModel,
    RunConfig,
    Runner,
    set_tracing_disabled,
)

set_tracing_disabled(True)

client = AsyncOpenAI(
    base_url="https://tokens.bd/v1",
    api_key=os.environ["TOKENS_API_KEY"],
)


class TokensModelProvider(ModelProvider):
    def get_model(self, model_name: str | None) -> Model:
        # The name arrives untouched, slash included.
        return OpenAIChatCompletionsModel(
            model=model_name or "deepseek/deepseek-v4.1-flash",
            openai_client=client,
        )


agent = Agent(
    name="Assistant",
    instructions="You are a helpful assistant.",
    model="deepseek/deepseek-v4.1-flash",
)


async def main() -> None:
    result = await Runner.run(
        agent,
        "Name two uses for asyncio.Semaphore.",
        run_config=RunConfig(model_provider=TokensModelProvider()),
    )
    print(result.final_output)


asyncio.run(main())
```

`RunConfig` applies to that one `Runner.run` call. A run without it goes to OpenAI's own endpoint with whatever `OPENAI_API_KEY` is set.

### Alternative: the global client

The SDK also has a global default, which is the shortest setup when every agent should use Tokens:

```python
import os
from openai import AsyncOpenAI
from agents import set_default_openai_api, set_default_openai_client, set_tracing_disabled

set_default_openai_client(
    AsyncOpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"]),
    use_for_tracing=False,
)
set_default_openai_api("chat_completions")
set_tracing_disabled(True)
```

With this setup, a plain string such as `model="deepseek/deepseek-v4.1-flash"` still goes through the SDK's default `MultiProvider`, which splits on the first slash. The part before the slash is not a provider the SDK knows, and the default behavior is to raise `UserError: Unknown prefix`. If the prefix is `openai`, the SDK silently sends only the part after the slash. To use model strings with the global setup, build the provider yourself and tell it to keep the full id:

```python
import os
from agents import MultiProvider, RunConfig

provider = MultiProvider(
    openai_base_url="https://tokens.bd/v1",
    openai_api_key=os.environ["TOKENS_API_KEY"],
    openai_use_responses=False,         # Chat Completions
    openai_prefix_mode="model_id",      # keep a leading "openai/" as part of the id
    unknown_prefix_mode="model_id",     # keep any other "provider/" as part of the id
)

run_config = RunConfig(model_provider=provider)
```

`openai_prefix_mode` and `unknown_prefix_mode` are documented by the SDK for sending namespaced ids such as `openrouter/openai/gpt-4.1-mini` to an OpenAI-compatible backend. The simpler routes (a model object, or the `ModelProvider` above) avoid the question, so prefer them.

:::note
`use_for_tracing=False` stops the SDK from using this client's key (your Tokens key) to upload traces to OpenAI. The default is `True`. The argument is described in the SDK's source docstring, not in its guide pages, so check it against the version you install.
:::

## Tracing: turn it off

The SDK uploads traces to OpenAI's servers by default. Without an OpenAI platform key you get 401 errors in your logs, even though the agent itself works. The documentation gives three ways to turn tracing off:

| Scope        | How                                                                      |
| ------------ | ------------------------------------------------------------------------ |
| Whole process | `set_tracing_disabled(True)`, or the environment variable `OPENAI_AGENTS_DISABLE_TRACING=1` |
| One run      | `RunConfig(tracing_disabled=True)`                                       |

You can also keep tracing and set a separate OpenAI key only for uploads with `set_tracing_export_api_key(...)`. That key has to come from platform.openai.com. It is not your Tokens key, and your prompts then go to OpenAI as trace data. If your prompts are private, leave tracing off.

## Tools

Give the agent Python functions with type hints and a docstring. The SDK builds the tool schema from them.

```python title="tools.py"
import asyncio
import os

from openai import AsyncOpenAI
from agents import (
    Agent,
    OpenAIChatCompletionsModel,
    Runner,
    function_tool,
    set_tracing_disabled,
)

set_tracing_disabled(True)

client = AsyncOpenAI(
    base_url="https://tokens.bd/v1",
    api_key=os.environ["TOKENS_API_KEY"],
)


@function_tool
def get_weather(city: str) -> str:
    """Current weather for a city.

    Args:
        city: The city to look up.
    """
    return f"{city}: light rain, 29C"  # replace with a real lookup


agent = Agent(
    name="Weather helper",
    instructions="Use the tool when the user asks about the weather.",
    model=OpenAIChatCompletionsModel(model="deepseek/deepseek-v4.1-flash", openai_client=client),
    tools=[get_weather],
)


async def main() -> None:
    result = await Runner.run(agent, "Do I need an umbrella in Dhaka?")
    print(result.final_output)


asyncio.run(main())
```

`function_tool` is the decorator in the released package. Newer SDK documentation imports the same decorator as `from agents.decorators import tool`, which is an alias for it. Either works on the version checked.

Each tool call is another request to Tokens, so a run with several tool rounds costs several requests and counts against your [rate limits](/docs/rate-limits). Tool calling needs a model that supports it ([Tool calling](/docs/tool-calling)).

The SDK's documentation lists tools that only work on the Responses API (for example `ToolSearchTool`), and says they are rejected on Chat Completions backends. Hosted tools such as web search or file search run on OpenAI's side and are not available through Tokens.

## Stream the output

`Runner.run_streamed` returns a result you iterate for events. This is the SDK's documented text-streaming loop:

```python title="stream.py"
import asyncio
import os

from openai import AsyncOpenAI
from openai.types.responses import ResponseTextDeltaEvent
from agents import Agent, OpenAIChatCompletionsModel, Runner, set_tracing_disabled

set_tracing_disabled(True)

client = AsyncOpenAI(
    base_url="https://tokens.bd/v1",
    api_key=os.environ["TOKENS_API_KEY"],
)

agent = Agent(
    name="Writer",
    instructions="You write short, plain answers.",
    model=OpenAIChatCompletionsModel(model="deepseek/deepseek-v4.1-flash", openai_client=client),
)


async def main() -> None:
    result = Runner.run_streamed(agent, input="Write a haiku about merge conflicts.")
    async for event in result.stream_events():
        if event.type == "raw_response_event" and isinstance(event.data, ResponseTextDeltaEvent):
            print(event.data.delta, end="", flush=True)
    print()


asyncio.run(main())
```

Keep reading the stream until it ends. The run is not finished until the iterator is. The SDK's streaming page shows this loop for the default setup. It does not say whether it behaves the same with `OpenAIChatCompletionsModel`, so test it once with your model. See also [Streaming](/docs/streaming).

Tokens streams carry token counts only when the request asks for them. On Chat Completions the SDK has a `ModelSettings(include_usage=True)` field for that (the field is documented as available for Chat Completions only). Import `ModelSettings` from `agents` and pass it as `Agent(..., model_settings=ModelSettings(include_usage=True))`.

If a provider sends broken tool-call fragments while streaming, the SDK has an option, `openai_buffer_streamed_tool_calls=True` on `MultiProvider`, that buffers them.

## TypeScript

The TypeScript package is `@openai/agents`. It needs `zod` 4 as a peer dependency, and the `openai` package (version 7.2 or newer) when you pass your own client.

```bash
npm install @openai/agents zod openai
```

```ts title="agent.ts"
import OpenAI from "openai";
import {
  Agent,
  OpenAIChatCompletionsModel,
  run,
  setTracingDisabled,
  tool,
} from "@openai/agents";
import { z } from "zod";

setTracingDisabled(true);

const client = new OpenAI({
  apiKey: process.env.TOKENS_API_KEY,
  baseURL: "https://tokens.bd/v1",
});

const getWeather = tool({
  name: "get_weather",
  description: "Current weather for a city",
  parameters: z.object({ city: z.string() }),
  async execute({ city }) {
    return `${city}: light rain, 29C`; // replace with a real lookup
  },
});

const agent = new Agent({
  name: "Weather helper",
  instructions: "Use the tool when the user asks about the weather.",
  model: new OpenAIChatCompletionsModel(client, "deepseek/deepseek-v4.1-flash"),
  tools: [getWeather],
});

const result = await run(agent, "Do I need an umbrella in Dhaka?");
console.log(result.finalOutput);
```

Passing an `OpenAIChatCompletionsModel` instance keeps that agent on Chat Completions, and the model id is not parsed. The SDK's own example also offers a global route, `setDefaultOpenAIClient(client)` with `setOpenAIAPI("chat_completions")`, and a `new Runner({ modelProvider })` route with an `OpenAIProvider({ openAIClient: client })`. The TypeScript documentation does not say how those routes treat a model name with a slash, so use the model object above.

To stream text in Node:

```ts
const stream = await run(agent, "Write a haiku about merge conflicts.", { stream: true });
stream.toTextStream({ compatibleWithNodeStreams: true }).pipe(process.stdout);
await stream.completed;
```

You can turn tracing off with `setTracingDisabled(true)` as above, or with `OPENAI_AGENTS_DISABLE_TRACING=1`. Run this code on a server or in a CLI, not in a browser bundle: Tokens does not send CORS headers and the key would be visible to anyone who opens dev tools.

## Check that it works

Run `python agent.py`. A short answer printed to the terminal means the key, base URL and model id are right. Then open Usage analytics in the [dashboard](/dashboard) and look for the request. If the call fails, test the endpoint with [cURL](/docs/curl) first: a 200 there with a failure in your agent points at the SDK setup, not the key.

## Choosing a model

Agents loop: they call tools, read results and call again, so the model must handle tool calls well and keep a long context. Some models answer plain chat but fail on tools. [Choosing a model](/docs/choosing-a-model) covers which suit agent work, and each model's page in [/models](/models) shows its context window and whether tools are supported. Set `max_tokens` through `ModelSettings` if you want to cap each reply. A run with several tool rounds can use a lot of tokens, so set a spend cap on the key you give an agent ([API keys](/docs/api-keys)).

## Limits and what does not work

- **Responses-only features.** Hosted tools (web search, file search, code interpreter), `previous_response_id` and other Responses-only fields are not available on the Chat Completions path. The SDK drops Responses-only fields silently unless you turn on strict validation (`strict_feature_validation=True` on `OpenAIProvider`, `openai_strict_feature_validation=True` on `MultiProvider`).
- **Structured output.** The SDK sends `json_schema` response formats. If the model behind a Tokens id does not support them, the upstream rejects the request with a 400 (`invalid_request`). Use a model that supports structured output ([Structured output](/docs/structured-output)).
- **Audio and Realtime.** The Chat Completions adapter raises `AgentsException("Audio is not currently supported")` for audio output. Realtime and voice agents need OpenAI's own endpoints and do not go through Tokens.
- **Tracing.** Traces go to OpenAI, not Tokens, and the Tokens key cannot upload them.
- **Empty replies on `finish_reason="length"`.** The adapter raises `ModelBehaviorError` when the model stops at the length limit with no output. That is a token or reasoning budget problem. Raise `max_tokens` or pick a model with less hidden reasoning.

## Troubleshooting

**`UserError: Unknown prefix: <first part of your model id>`.** The model was given to the agent as a string and the SDK's default provider read the part before the slash as a provider prefix. Pass an `OpenAIChatCompletionsModel` object, use the `ModelProvider` above, or set `unknown_prefix_mode="model_id"` on a `MultiProvider`.

**404 `model_not_found`.** Either the id has a typo, or the SDK removed a leading `openai/` from it. Compare the id in your Tokens usage log with the one you meant to send, and check it against `GET https://tokens.bd/v1/models`.

**404 or 400 from `/v1/responses`.** The SDK is still on the Responses API. Use `OpenAIChatCompletionsModel` or `set_default_openai_api("chat_completions")`.

**401 errors from `api.openai.com` or "incorrect API key" in the log, and the agent still answers.** These are trace uploads. Turn tracing off.

**401 `invalid_api_key` from Tokens.** The key in `TOKENS_API_KEY` is wrong or revoked. Create a new one in [/dashboard/keys](/dashboard/keys).

**402 `insufficient_credits`, 403 `model_not_allowed_on_key`, 429 `window_exhausted`.** These are account limits, not SDK problems. See [Errors](/docs/errors) and [Troubleshooting](/docs/troubleshooting). An agent that fires many requests in parallel can hit `429 concurrency_limit`: lower the parallelism ([Rate limits](/docs/rate-limits)).

**Errors surface as SDK exceptions.** The SDK wraps the OpenAI client, so a Tokens error arrives as an `openai` exception. Read `e.code` and `e.response.headers.get("x-tokens-request-id")` and include the request id when you contact [support](/docs/support).

---
Page: https://tokens.bd/docs/openai-agents-sdk
