POST https://tokens.bd/v1/responses passes through the OpenAI Responses API. It exists mainly for clients built on it, Codex CLI being the obvious one. If you are writing new code and have no reason to prefer it, chat completions is the more widely supported choice across upstream models.
When to use the Responses API#
Use /v1/responses when | Use /v1/chat/completions when |
|---|---|
Your tool only speaks Responses (Codex CLI with wire_api = "responses") | You want the broadest model compatibility |
You already have code written against client.responses.create | You use frameworks that expect chat completions |
| You want typed output items and event names | You need n > 1 or other chat-only fields |
Support for this endpoint depends on the upstream that serves the model. The gateway forwards the request as is; if the provider behind a model doesn't implement Responses, the call fails with a 400 or 404 from upstream. When that happens, use chat completions for that model.
Responses API request example#
curl https://tokens.bd/v1/responses \
-H "Authorization: Bearer $TOKENS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek/deepseek-v4.1-flash",
"instructions": "You are a concise senior engineer.",
"input": "Give me one reason to pin dependency versions.",
"max_output_tokens": 200
}'import os
from openai import OpenAI
client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"])
resp = client.responses.create(
model="deepseek/deepseek-v4.1-flash",
instructions="You are a concise senior engineer.",
input="Give me one reason to pin dependency versions.",
max_output_tokens=200,
)
print(resp.output_text)
print(resp.usage)A trimmed response:
{
"id": "resp_abc123",
"object": "response",
"status": "completed",
"model": "deepseek/deepseek-v4.1-flash",
"output": [
{
"type": "message",
"role": "assistant",
"content": [
{
"type": "output_text",
"text": "Reproducible builds: the same commit installs the same code everywhere."
}
]
}
],
"usage": {
"input_tokens": 24,
"output_tokens": 15,
"total_tokens": 39
}
}Fields commonly used in the request:
| Field | Notes |
|---|---|
model | Required. Any id from GET /v1/models. |
input | A string, or an array of input items (messages, tool outputs). |
instructions | System-style instructions. |
max_output_tokens | Output limit. The gateway uses it for the cost reservation and, on a low balance, may lower it (minimum 16). |
tools, tool_choice | Function tools, where the model supports them. Hosted tools such as web search or file search are provider features; don't assume they work through the gateway. |
stream | true for Server-Sent Events. |
Conversation state#
Tokens doesn't store prompts or responses, so do not count on previous_response_id or store: true to keep state between calls. Whether they work depends on the upstream, and a retry or failover can land on a different source that has never seen the earlier response. Send the full conversation in input each time; that works everywhere.
Stream Responses API events#
With "stream": true, the response is Server-Sent Events with typed events. The ones most clients care about:
| Event | Contains |
|---|---|
response.created | The response object, status in_progress |
response.output_text.delta | A text fragment in delta |
response.function_call_arguments.delta | A fragment of tool-call arguments |
response.completed | The final response object, including usage |
Unlike chat completions, you don't need stream_options to get usage: it arrives in response.completed.
stream = client.responses.create(
model="deepseek/deepseek-v4.1-flash",
input="Write a haiku about cache invalidation.",
stream=True,
)
for event in stream:
if event.type == "response.output_text.delta":
print(event.delta, end="", flush=True)
elif event.type == "response.completed":
print("\n", event.response.usage)Disconnect handling and timeouts work the same as for the other endpoints; see streaming.
Codex CLI uses /v1/responses#
Codex CLI talks to Tokens through this endpoint. The provider block in ~/.codex/config.toml looks like this, with the key read from TOKENS_API_KEY:
model = "deepseek/deepseek-v4.1-flash"
model_provider = "tokens"
[model_providers.tokens]
name = "Tokens"
base_url = "https://tokens.bd/v1"
env_key = "TOKENS_API_KEY"
wire_api = "responses"The Tokens CLI can write this for you; the quickstart has the one-line setup. If Codex reports errors for a specific model, check whether that model works on /v1/responses with the curl example above before digging into Codex settings.
Errors follow the shape described in errors, and limits are in rate limits.