Set "stream": true on chat completions, messages, responses or legacy completions and the gateway streams Server-Sent Events (SSE) from the upstream provider to you as they arrive. The gateway doesn't buffer or rewrite the content; it reads token counts off the stream for billing and passes the bytes through.
SSE format for chat completions#
Each event is a data: line with a JSON chunk, separated by a blank line. The stream ends with data: [DONE].
data: {"id":"chatcmpl-a1b2","object":"chat.completion.chunk","created":1790000000,"model":"deepseek/deepseek-v4.1-flash","choices":[{"index":0,"delta":{"role":"assistant","content":""},"finish_reason":null}]}
data: {"id":"chatcmpl-a1b2","object":"chat.completion.chunk","created":1790000000,"model":"deepseek/deepseek-v4.1-flash","choices":[{"index":0,"delta":{"content":"Retry"},"finish_reason":null}]}
data: {"id":"chatcmpl-a1b2","object":"chat.completion.chunk","created":1790000000,"model":"deepseek/deepseek-v4.1-flash","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}
data: {"id":"chatcmpl-a1b2","object":"chat.completion.chunk","created":1790000000,"model":"deepseek/deepseek-v4.1-flash","choices":[],"usage":{"prompt_tokens":18,"completion_tokens":42,"total_tokens":60}}
data: [DONE]Reasoning models may also send delta.reasoning_content, and tool calls arrive as delta.tool_calls fragments (see tool calling).
Get token usage while streaming with include_usage#
The chunk with "choices": [] and a usage object only appears if you ask for it:
{
"stream": true,
"stream_options": { "include_usage": true }
}Without it, you get no usage chunk, matching OpenAI's behavior. Billing doesn't depend on this flag: the gateway always meters the stream. The flag only controls whether you see the numbers.
If you set it, guard against the empty choices array in your loop. Code that does chunk.choices[0] unconditionally will crash on the last chunk.
SSE format for Anthropic messages#
/v1/messages streams named events: message_start, content_block_start, content_block_delta, content_block_stop, message_delta and message_stop. Input token usage is in message_start, output usage in message_delta, so there is no flag to set. The full sequence is shown in messages. The Responses API has its own event names, listed in responses.
Stream in Python and Node.js#
import os
from openai import OpenAI
client = OpenAI(
base_url="https://tokens.bd/v1",
api_key=os.environ["TOKENS_API_KEY"],
timeout=600, # seconds; long reasoning answers can take minutes
)
stream = client.chat.completions.create(
model="deepseek/deepseek-v4.1-flash",
messages=[{"role": "user", "content": "Explain exponential backoff with jitter."}],
stream=True,
stream_options={"include_usage": True},
)
finished = False
for chunk in stream:
if chunk.choices:
choice = chunk.choices[0]
if choice.delta.content:
print(choice.delta.content, end="", flush=True)
if choice.finish_reason:
finished = True
if chunk.usage:
print("\n", chunk.usage)
if not finished:
print("\n[stream ended without a finish_reason: treat the answer as incomplete]")import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://tokens.bd/v1",
apiKey: process.env.TOKENS_API_KEY,
timeout: 600_000, // ms
});
const controller = new AbortController();
// Cancel after 2 minutes, or wire this to a user's "stop" button.
const timer = setTimeout(() => controller.abort(), 120_000);
const stream = await client.chat.completions.create(
{
model: "deepseek/deepseek-v4.1-flash",
messages: [{ role: "user", content: "Explain exponential backoff with jitter." }],
stream: true,
stream_options: { include_usage: true },
},
{ signal: controller.signal }
);
try {
for await (const chunk of stream) {
const delta = chunk.choices[0]?.delta?.content;
if (delta) process.stdout.write(delta);
if (chunk.usage) console.log("\n", chunk.usage);
}
} finally {
clearTimeout(timer);
}For raw HTTP, curl -N disables output buffering so you can watch events arrive:
curl -N https://tokens.bd/v1/chat/completions \
-H "Authorization: Bearer $TOKENS_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"deepseek/deepseek-v4.1-flash","stream":true,"messages":[{"role":"user","content":"Count to five."}]}'Errors before and during a stream#
Errors that happen before the first byte, such as a bad key, no balance, rate limits or an upstream that refused the request, come back as a normal JSON error with the matching HTTP status, not as an SSE event. Check the status code before you start parsing events. The errors page lists every code.
Once the stream has started, the status is already 200. If the upstream fails midway, the connection closes without [DONE] (or without message_stop on messages). Treat a stream that ends without a finish_reason or stop event as incomplete.
The x-tokens-request-id response header arrives with the headers, before any content. Log it at the start of each request so you have it if the stream dies later.
Timeouts for long requests#
Long requests are fine. The gateway waits up to 600 seconds for an upstream to start responding, which covers reasoning models that think for minutes before the first token, and up to 300 seconds between chunks once a stream is flowing. If the upstream doesn't start in time, you get 504 upstream_timeout.
Your client's timeout needs to be at least as generous. Set it explicitly, as in the examples above, rather than relying on library defaults or a proxy in front of your app that cuts idle connections after 30 or 60 seconds.
Handle client disconnects#
If your client disconnects or aborts mid-stream, the gateway cancels the upstream request and bills for what was generated up to that point: the input plus the output already streamed. Nothing further is charged.
Before retrying a broken stream, remember that a retry sends and bills the full prompt again. For long agent prompts, that adds up. Retry streams that failed with a retryable error before any content arrived; for streams that broke partway, decide whether the partial output is usable first. Retry rules by error code are in errors.
Each open stream counts toward your account's concurrency limit until it finishes. Streams you abandon without closing hold a slot, so close or abort them explicitly. See rate limits.