Skip to content

Streaming

How Server-Sent Events streaming works through the Tokens gateway for chat completions and messages: the event format, usage chunks, timeouts, and handling disconnects.

On this page

Set "stream": true on chat completions, messages, responses or legacy completions and the gateway streams Server-Sent Events (SSE) from the upstream provider to you as they arrive. The gateway doesn't buffer or rewrite the content; it reads token counts off the stream for billing and passes the bytes through.

SSE format for chat completions#

Each event is a data: line with a JSON chunk, separated by a blank line. The stream ends with data: [DONE].

text
data: {"id":"chatcmpl-a1b2","object":"chat.completion.chunk","created":1790000000,"model":"deepseek/deepseek-v4.1-flash","choices":[{"index":0,"delta":{"role":"assistant","content":""},"finish_reason":null}]}

data: {"id":"chatcmpl-a1b2","object":"chat.completion.chunk","created":1790000000,"model":"deepseek/deepseek-v4.1-flash","choices":[{"index":0,"delta":{"content":"Retry"},"finish_reason":null}]}

data: {"id":"chatcmpl-a1b2","object":"chat.completion.chunk","created":1790000000,"model":"deepseek/deepseek-v4.1-flash","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}

data: {"id":"chatcmpl-a1b2","object":"chat.completion.chunk","created":1790000000,"model":"deepseek/deepseek-v4.1-flash","choices":[],"usage":{"prompt_tokens":18,"completion_tokens":42,"total_tokens":60}}

data: [DONE]

Reasoning models may also send delta.reasoning_content, and tool calls arrive as delta.tool_calls fragments (see tool calling).

Get token usage while streaming with include_usage#

The chunk with "choices": [] and a usage object only appears if you ask for it:

json
{
  "stream": true,
  "stream_options": { "include_usage": true }
}

Without it, you get no usage chunk, matching OpenAI's behavior. Billing doesn't depend on this flag: the gateway always meters the stream. The flag only controls whether you see the numbers.

If you set it, guard against the empty choices array in your loop. Code that does chunk.choices[0] unconditionally will crash on the last chunk.

SSE format for Anthropic messages#

/v1/messages streams named events: message_start, content_block_start, content_block_delta, content_block_stop, message_delta and message_stop. Input token usage is in message_start, output usage in message_delta, so there is no flag to set. The full sequence is shown in messages. The Responses API has its own event names, listed in responses.

Stream in Python and Node.js#

import os
from openai import OpenAI

client = OpenAI(
    base_url="https://tokens.bd/v1",
    api_key=os.environ["TOKENS_API_KEY"],
    timeout=600,  # seconds; long reasoning answers can take minutes
)

stream = client.chat.completions.create(
    model="deepseek/deepseek-v4.1-flash",
    messages=[{"role": "user", "content": "Explain exponential backoff with jitter."}],
    stream=True,
    stream_options={"include_usage": True},
)

finished = False
for chunk in stream:
    if chunk.choices:
        choice = chunk.choices[0]
        if choice.delta.content:
            print(choice.delta.content, end="", flush=True)
        if choice.finish_reason:
            finished = True
    if chunk.usage:
        print("\n", chunk.usage)

if not finished:
    print("\n[stream ended without a finish_reason: treat the answer as incomplete]")

For raw HTTP, curl -N disables output buffering so you can watch events arrive:

bash
curl -N https://tokens.bd/v1/chat/completions \
  -H "Authorization: Bearer $TOKENS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"deepseek/deepseek-v4.1-flash","stream":true,"messages":[{"role":"user","content":"Count to five."}]}'

Errors before and during a stream#

Errors that happen before the first byte, such as a bad key, no balance, rate limits or an upstream that refused the request, come back as a normal JSON error with the matching HTTP status, not as an SSE event. Check the status code before you start parsing events. The errors page lists every code.

Once the stream has started, the status is already 200. If the upstream fails midway, the connection closes without [DONE] (or without message_stop on messages). Treat a stream that ends without a finish_reason or stop event as incomplete.

The x-tokens-request-id response header arrives with the headers, before any content. Log it at the start of each request so you have it if the stream dies later.

Timeouts for long requests#

Long requests are fine. The gateway waits up to 600 seconds for an upstream to start responding, which covers reasoning models that think for minutes before the first token, and up to 300 seconds between chunks once a stream is flowing. If the upstream doesn't start in time, you get 504 upstream_timeout.

Your client's timeout needs to be at least as generous. Set it explicitly, as in the examples above, rather than relying on library defaults or a proxy in front of your app that cuts idle connections after 30 or 60 seconds.

Handle client disconnects#

If your client disconnects or aborts mid-stream, the gateway cancels the upstream request and bills for what was generated up to that point: the input plus the output already streamed. Nothing further is charged.

Before retrying a broken stream, remember that a retry sends and bills the full prompt again. For long agent prompts, that adds up. Retry streams that failed with a retryable error before any content arrived; for streams that broke partway, decide whether the partial output is usable first. Retry rules by error code are in errors.

Each open stream counts toward your account's concurrency limit until it finishes. Streams you abandon without closing hold a slot, so close or abort them explicitly. See rate limits.

Was this page helpful?

Still stuck? Open a support ticket

Need help configuring your agent?

Test your connection with the connection tester, or create an API key.