Skip to content

Responses API

POST /v1/responses, the OpenAI Responses API: when to use it instead of chat completions, request and response examples, max_output_tokens and streaming events.

On this page

POST https://tokens.bd/v1/responses passes through the OpenAI Responses API. It exists mainly for clients built on it, Codex CLI being the obvious one. If you are writing new code and have no reason to prefer it, chat completions is the more widely supported choice across upstream models.

When to use the Responses API#

Use /v1/responses whenUse /v1/chat/completions when
Your tool only speaks Responses (Codex CLI with wire_api = "responses")You want the broadest model compatibility
You already have code written against client.responses.createYou use frameworks that expect chat completions
You want typed output items and event namesYou need n > 1 or other chat-only fields

Support for this endpoint depends on the upstream that serves the model. The gateway forwards the request as is; if the provider behind a model doesn't implement Responses, the call fails with a 400 or 404 from upstream. When that happens, use chat completions for that model.

Responses API request example#

curl https://tokens.bd/v1/responses \
  -H "Authorization: Bearer $TOKENS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek/deepseek-v4.1-flash",
    "instructions": "You are a concise senior engineer.",
    "input": "Give me one reason to pin dependency versions.",
    "max_output_tokens": 200
  }'

A trimmed response:

json
{
  "id": "resp_abc123",
  "object": "response",
  "status": "completed",
  "model": "deepseek/deepseek-v4.1-flash",
  "output": [
    {
      "type": "message",
      "role": "assistant",
      "content": [
        {
          "type": "output_text",
          "text": "Reproducible builds: the same commit installs the same code everywhere."
        }
      ]
    }
  ],
  "usage": {
    "input_tokens": 24,
    "output_tokens": 15,
    "total_tokens": 39
  }
}

Fields commonly used in the request:

FieldNotes
modelRequired. Any id from GET /v1/models.
inputA string, or an array of input items (messages, tool outputs).
instructionsSystem-style instructions.
max_output_tokensOutput limit. The gateway uses it for the cost reservation and, on a low balance, may lower it (minimum 16).
tools, tool_choiceFunction tools, where the model supports them. Hosted tools such as web search or file search are provider features; don't assume they work through the gateway.
streamtrue for Server-Sent Events.

Conversation state#

Tokens doesn't store prompts or responses, so do not count on previous_response_id or store: true to keep state between calls. Whether they work depends on the upstream, and a retry or failover can land on a different source that has never seen the earlier response. Send the full conversation in input each time; that works everywhere.

Stream Responses API events#

With "stream": true, the response is Server-Sent Events with typed events. The ones most clients care about:

EventContains
response.createdThe response object, status in_progress
response.output_text.deltaA text fragment in delta
response.function_call_arguments.deltaA fragment of tool-call arguments
response.completedThe final response object, including usage

Unlike chat completions, you don't need stream_options to get usage: it arrives in response.completed.

python
stream = client.responses.create(
    model="deepseek/deepseek-v4.1-flash",
    input="Write a haiku about cache invalidation.",
    stream=True,
)
for event in stream:
    if event.type == "response.output_text.delta":
        print(event.delta, end="", flush=True)
    elif event.type == "response.completed":
        print("\n", event.response.usage)

Disconnect handling and timeouts work the same as for the other endpoints; see streaming.

Codex CLI uses /v1/responses#

Codex CLI talks to Tokens through this endpoint. The provider block in ~/.codex/config.toml looks like this, with the key read from TOKENS_API_KEY:

/.codex/config.toml
model = "deepseek/deepseek-v4.1-flash"
model_provider = "tokens"

[model_providers.tokens]
name = "Tokens"
base_url = "https://tokens.bd/v1"
env_key = "TOKENS_API_KEY"
wire_api = "responses"

The Tokens CLI can write this for you; the quickstart has the one-line setup. If Codex reports errors for a specific model, check whether that model works on /v1/responses with the curl example above before digging into Codex settings.

Errors follow the shape described in errors, and limits are in rate limits.

Was this page helpful?

Still stuck? Open a support ticket

Need help configuring your agent?

Test your connection with the connection tester, or create an API key.