Skip to content

OpenAI Agents SDK

Run agents built with the OpenAI Agents SDK (Python and TypeScript) on Tokens: Chat Completions model class, the slash-in-model-id problem, tracing, tools and streaming.

Works withOpenAI Agents SDK
On this page

The OpenAI Agents SDK is OpenAI's library for building agents: an Agent with instructions and tools, run by a Runner. It talks to models through the OpenAI API, so it can reach Tokens at https://tokens.bd/v1. Two defaults get in the way, and both are easy to fix: the SDK calls the Responses API unless you tell it otherwise, and it reads a model id that contains a slash as provider/model. Every Tokens model id contains a slash, so this page shows how to set both up.

Checked against the documentation

Based on the OpenAI Agents SDK documentation, checked October 2026 against the Python package openai-agents 0.23.1 (released 2 October 2026) and the TypeScript package @openai/agents 0.20.0. The code was checked against the documentation and the SDK's own examples, not run end to end against Tokens. The SDK changes quickly, so compare with the official docs if something here no longer matches your version.

What you need#

  • A Tokens key from API keys, exported as TOKENS_API_KEY.
  • A model id from /models, for example deepseek/deepseek-v4.1-flash. Agents call tools, so pick a model that supports tool calling (Choosing a model).
  • Python 3.10 or newer for the Python package. The TypeScript package needs the openai package version 7.2 or newer when you pass your own client.
bash
export TOKENS_API_KEY="tok_live_your_key"

Why the defaults need changing#

Default in the SDKWhat happens with TokensFix
Uses the Responses APITokens passes /v1/responses through, but it only works if the provider behind the model implements it (Responses). Chat Completions works for every model.Use OpenAIChatCompletionsModel, or set_default_openai_api("chat_completions").
Reads prefix/name as a provider prefixThe SDK documents that unknown prefixes raise UserError, and that openai/... is shortened to the part after the slash. A Tokens id such as deepseek/deepseek-v4.1-flash hits this rule.Pass a model object, use a custom ModelProvider, or switch the prefix modes.
Uploads traces to OpenAITraces go to OpenAI's servers and need an OpenAI key. With a Tokens key the upload fails with a 401 in your logs.Turn tracing off.

The SDK's documentation lists these under its "non-OpenAI models" section. It recommends Chat Completions when a provider lacks Responses support, and says unknown prefixes raise UserError instead of being passed through.

Python: use a Chat Completions model object#

This is the setup to start with. You create an AsyncOpenAI client that points at Tokens, wrap it in OpenAIChatCompletionsModel, and give that object to the agent. Because the agent receives a model object and not a model name string, the SDK does not parse the model id, so the slash is not an issue.

bash
pip install --upgrade openai-agents
agent.py
import asyncio
import os

from openai import AsyncOpenAI
from agents import Agent, OpenAIChatCompletionsModel, Runner, set_tracing_disabled

set_tracing_disabled(True)  # traces would go to OpenAI; see the Tracing section

client = AsyncOpenAI(
    base_url="https://tokens.bd/v1",
    api_key=os.environ["TOKENS_API_KEY"],
)

agent = Agent(
    name="Reviewer",
    instructions="You are a concise senior engineer. Answer in at most three sentences.",
    model=OpenAIChatCompletionsModel(model="deepseek/deepseek-v4.1-flash", openai_client=client),
)


async def main() -> None:
    result = await Runner.run(agent, "When should I use a dataclass instead of a dict?")
    print(result.final_output)


asyncio.run(main())

base_url keeps the /v1 at the end. The model string goes to Tokens exactly as written, so it must match an id from /models or GET https://tokens.bd/v1/models.

Python: set a model for every agent in one place#

If you have many agents, a custom ModelProvider passed through RunConfig applies one model setup to a whole run, and your Agent objects can keep plain model names. This follows the SDK's own custom_example_provider.py example.

provider.py
import asyncio
import os

from openai import AsyncOpenAI
from agents import (
    Agent,
    Model,
    ModelProvider,
    OpenAIChatCompletionsModel,
    RunConfig,
    Runner,
    set_tracing_disabled,
)

set_tracing_disabled(True)

client = AsyncOpenAI(
    base_url="https://tokens.bd/v1",
    api_key=os.environ["TOKENS_API_KEY"],
)


class TokensModelProvider(ModelProvider):
    def get_model(self, model_name: str | None) -> Model:
        # The name arrives untouched, slash included.
        return OpenAIChatCompletionsModel(
            model=model_name or "deepseek/deepseek-v4.1-flash",
            openai_client=client,
        )


agent = Agent(
    name="Assistant",
    instructions="You are a helpful assistant.",
    model="deepseek/deepseek-v4.1-flash",
)


async def main() -> None:
    result = await Runner.run(
        agent,
        "Name two uses for asyncio.Semaphore.",
        run_config=RunConfig(model_provider=TokensModelProvider()),
    )
    print(result.final_output)


asyncio.run(main())

RunConfig applies to that one Runner.run call. A run without it goes to OpenAI's own endpoint with whatever OPENAI_API_KEY is set.

Alternative: the global client#

The SDK also has a global default, which is the shortest setup when every agent should use Tokens:

python
import os
from openai import AsyncOpenAI
from agents import set_default_openai_api, set_default_openai_client, set_tracing_disabled

set_default_openai_client(
    AsyncOpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"]),
    use_for_tracing=False,
)
set_default_openai_api("chat_completions")
set_tracing_disabled(True)

With this setup, a plain string such as model="deepseek/deepseek-v4.1-flash" still goes through the SDK's default MultiProvider, which splits on the first slash. The part before the slash is not a provider the SDK knows, and the default behavior is to raise UserError: Unknown prefix. If the prefix is openai, the SDK silently sends only the part after the slash. To use model strings with the global setup, build the provider yourself and tell it to keep the full id:

python
import os
from agents import MultiProvider, RunConfig

provider = MultiProvider(
    openai_base_url="https://tokens.bd/v1",
    openai_api_key=os.environ["TOKENS_API_KEY"],
    openai_use_responses=False,         # Chat Completions
    openai_prefix_mode="model_id",      # keep a leading "openai/" as part of the id
    unknown_prefix_mode="model_id",     # keep any other "provider/" as part of the id
)

run_config = RunConfig(model_provider=provider)

openai_prefix_mode and unknown_prefix_mode are documented by the SDK for sending namespaced ids such as openrouter/openai/gpt-4.1-mini to an OpenAI-compatible backend. The simpler routes (a model object, or the ModelProvider above) avoid the question, so prefer them.

Note

use_for_tracing=False stops the SDK from using this client's key (your Tokens key) to upload traces to OpenAI. The default is True. The argument is described in the SDK's source docstring, not in its guide pages, so check it against the version you install.

Tracing: turn it off#

The SDK uploads traces to OpenAI's servers by default. Without an OpenAI platform key you get 401 errors in your logs, even though the agent itself works. The documentation gives three ways to turn tracing off:

ScopeHow
Whole processset_tracing_disabled(True), or the environment variable OPENAI_AGENTS_DISABLE_TRACING=1
One runRunConfig(tracing_disabled=True)

You can also keep tracing and set a separate OpenAI key only for uploads with set_tracing_export_api_key(...). That key has to come from platform.openai.com. It is not your Tokens key, and your prompts then go to OpenAI as trace data. If your prompts are private, leave tracing off.

Tools#

Give the agent Python functions with type hints and a docstring. The SDK builds the tool schema from them.

tools.py
import asyncio
import os

from openai import AsyncOpenAI
from agents import (
    Agent,
    OpenAIChatCompletionsModel,
    Runner,
    function_tool,
    set_tracing_disabled,
)

set_tracing_disabled(True)

client = AsyncOpenAI(
    base_url="https://tokens.bd/v1",
    api_key=os.environ["TOKENS_API_KEY"],
)


@function_tool
def get_weather(city: str) -> str:
    """Current weather for a city.

    Args:
        city: The city to look up.
    """
    return f"{city}: light rain, 29C"  # replace with a real lookup


agent = Agent(
    name="Weather helper",
    instructions="Use the tool when the user asks about the weather.",
    model=OpenAIChatCompletionsModel(model="deepseek/deepseek-v4.1-flash", openai_client=client),
    tools=[get_weather],
)


async def main() -> None:
    result = await Runner.run(agent, "Do I need an umbrella in Dhaka?")
    print(result.final_output)


asyncio.run(main())

function_tool is the decorator in the released package. Newer SDK documentation imports the same decorator as from agents.decorators import tool, which is an alias for it. Either works on the version checked.

Each tool call is another request to Tokens, so a run with several tool rounds costs several requests and counts against your rate limits. Tool calling needs a model that supports it (Tool calling).

The SDK's documentation lists tools that only work on the Responses API (for example ToolSearchTool), and says they are rejected on Chat Completions backends. Hosted tools such as web search or file search run on OpenAI's side and are not available through Tokens.

Stream the output#

Runner.run_streamed returns a result you iterate for events. This is the SDK's documented text-streaming loop:

stream.py
import asyncio
import os

from openai import AsyncOpenAI
from openai.types.responses import ResponseTextDeltaEvent
from agents import Agent, OpenAIChatCompletionsModel, Runner, set_tracing_disabled

set_tracing_disabled(True)

client = AsyncOpenAI(
    base_url="https://tokens.bd/v1",
    api_key=os.environ["TOKENS_API_KEY"],
)

agent = Agent(
    name="Writer",
    instructions="You write short, plain answers.",
    model=OpenAIChatCompletionsModel(model="deepseek/deepseek-v4.1-flash", openai_client=client),
)


async def main() -> None:
    result = Runner.run_streamed(agent, input="Write a haiku about merge conflicts.")
    async for event in result.stream_events():
        if event.type == "raw_response_event" and isinstance(event.data, ResponseTextDeltaEvent):
            print(event.data.delta, end="", flush=True)
    print()


asyncio.run(main())

Keep reading the stream until it ends. The run is not finished until the iterator is. The SDK's streaming page shows this loop for the default setup. It does not say whether it behaves the same with OpenAIChatCompletionsModel, so test it once with your model. See also Streaming.

Tokens streams carry token counts only when the request asks for them. On Chat Completions the SDK has a ModelSettings(include_usage=True) field for that (the field is documented as available for Chat Completions only). Import ModelSettings from agents and pass it as Agent(..., model_settings=ModelSettings(include_usage=True)).

If a provider sends broken tool-call fragments while streaming, the SDK has an option, openai_buffer_streamed_tool_calls=True on MultiProvider, that buffers them.

TypeScript#

The TypeScript package is @openai/agents. It needs zod 4 as a peer dependency, and the openai package (version 7.2 or newer) when you pass your own client.

bash
npm install @openai/agents zod openai
agent.ts
import OpenAI from "openai";
import {
  Agent,
  OpenAIChatCompletionsModel,
  run,
  setTracingDisabled,
  tool,
} from "@openai/agents";
import { z } from "zod";

setTracingDisabled(true);

const client = new OpenAI({
  apiKey: process.env.TOKENS_API_KEY,
  baseURL: "https://tokens.bd/v1",
});

const getWeather = tool({
  name: "get_weather",
  description: "Current weather for a city",
  parameters: z.object({ city: z.string() }),
  async execute({ city }) {
    return `${city}: light rain, 29C`; // replace with a real lookup
  },
});

const agent = new Agent({
  name: "Weather helper",
  instructions: "Use the tool when the user asks about the weather.",
  model: new OpenAIChatCompletionsModel(client, "deepseek/deepseek-v4.1-flash"),
  tools: [getWeather],
});

const result = await run(agent, "Do I need an umbrella in Dhaka?");
console.log(result.finalOutput);

Passing an OpenAIChatCompletionsModel instance keeps that agent on Chat Completions, and the model id is not parsed. The SDK's own example also offers a global route, setDefaultOpenAIClient(client) with setOpenAIAPI("chat_completions"), and a new Runner({ modelProvider }) route with an OpenAIProvider({ openAIClient: client }). The TypeScript documentation does not say how those routes treat a model name with a slash, so use the model object above.

To stream text in Node:

ts
const stream = await run(agent, "Write a haiku about merge conflicts.", { stream: true });
stream.toTextStream({ compatibleWithNodeStreams: true }).pipe(process.stdout);
await stream.completed;

You can turn tracing off with setTracingDisabled(true) as above, or with OPENAI_AGENTS_DISABLE_TRACING=1. Run this code on a server or in a CLI, not in a browser bundle: Tokens does not send CORS headers and the key would be visible to anyone who opens dev tools.

Check that it works#

Run python agent.py. A short answer printed to the terminal means the key, base URL and model id are right. Then open Usage analytics in the dashboard and look for the request. If the call fails, test the endpoint with cURL first: a 200 there with a failure in your agent points at the SDK setup, not the key.

Choosing a model#

Agents loop: they call tools, read results and call again, so the model must handle tool calls well and keep a long context. Some models answer plain chat but fail on tools. Choosing a model covers which suit agent work, and each model's page in /models shows its context window and whether tools are supported. Set max_tokens through ModelSettings if you want to cap each reply. A run with several tool rounds can use a lot of tokens, so set a spend cap on the key you give an agent (API keys).

Limits and what does not work#

  • Responses-only features. Hosted tools (web search, file search, code interpreter), previous_response_id and other Responses-only fields are not available on the Chat Completions path. The SDK drops Responses-only fields silently unless you turn on strict validation (strict_feature_validation=True on OpenAIProvider, openai_strict_feature_validation=True on MultiProvider).
  • Structured output. The SDK sends json_schema response formats. If the model behind a Tokens id does not support them, the upstream rejects the request with a 400 (invalid_request). Use a model that supports structured output (Structured output).
  • Audio and Realtime. The Chat Completions adapter raises AgentsException("Audio is not currently supported") for audio output. Realtime and voice agents need OpenAI's own endpoints and do not go through Tokens.
  • Tracing. Traces go to OpenAI, not Tokens, and the Tokens key cannot upload them.
  • Empty replies on finish_reason="length". The adapter raises ModelBehaviorError when the model stops at the length limit with no output. That is a token or reasoning budget problem. Raise max_tokens or pick a model with less hidden reasoning.

Troubleshooting#

UserError: Unknown prefix: <first part of your model id>. The model was given to the agent as a string and the SDK's default provider read the part before the slash as a provider prefix. Pass an OpenAIChatCompletionsModel object, use the ModelProvider above, or set unknown_prefix_mode="model_id" on a MultiProvider.

404 model_not_found. Either the id has a typo, or the SDK removed a leading openai/ from it. Compare the id in your Tokens usage log with the one you meant to send, and check it against GET https://tokens.bd/v1/models.

404 or 400 from /v1/responses. The SDK is still on the Responses API. Use OpenAIChatCompletionsModel or set_default_openai_api("chat_completions").

401 errors from api.openai.com or "incorrect API key" in the log, and the agent still answers. These are trace uploads. Turn tracing off.

401 invalid_api_key from Tokens. The key in TOKENS_API_KEY is wrong or revoked. Create a new one in /dashboard/keys.

402 insufficient_credits, 403 model_not_allowed_on_key, 429 window_exhausted. These are account limits, not SDK problems. See Errors and Troubleshooting. An agent that fires many requests in parallel can hit 429 concurrency_limit: lower the parallelism (Rate limits).

Errors surface as SDK exceptions. The SDK wraps the OpenAI client, so a Tokens error arrives as an openai exception. Read e.code and e.response.headers.get("x-tokens-request-id") and include the request id when you contact support.

Was this page helpful?

Still stuck? Open a support ticket

Need help configuring your agent?

Test your connection with the connection tester, or create an API key.