POST https://tokens.bd/v1/chat/completions is the OpenAI-compatible chat completions API, and the endpoint most SDKs and coding agents use. The request and response follow OpenAI's format; the gateway checks your key, plan and limits, then forwards the body to the upstream provider for the model you named.
Make a chat completions request#
curl https://tokens.bd/v1/chat/completions \
-H "Authorization: Bearer $TOKENS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek/deepseek-v4.1-flash",
"messages": [
{"role": "system", "content": "You are a concise senior engineer."},
{"role": "user", "content": "What does HTTP 429 mean?"}
],
"temperature": 0.2,
"max_tokens": 300
}'import os
from openai import OpenAI
client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"])
resp = client.chat.completions.create(
model="deepseek/deepseek-v4.1-flash",
messages=[
{"role": "system", "content": "You are a concise senior engineer."},
{"role": "user", "content": "What does HTTP 429 mean?"},
],
temperature=0.2,
max_tokens=300,
)
print(resp.choices[0].message.content)
print(resp.usage)import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://tokens.bd/v1", apiKey: process.env.TOKENS_API_KEY });
const resp = await client.chat.completions.create({
model: "deepseek/deepseek-v4.1-flash",
messages: [
{ role: "system", content: "You are a concise senior engineer." },
{ role: "user", content: "What does HTTP 429 mean?" },
],
temperature: 0.2,
max_tokens: 300,
});
console.log(resp.choices[0].message.content, resp.usage);The Content-Type: application/json header matters. Without it the gateway can't read the body, and you get a 400 saying the request must specify a model field even though it does.
Request fields#
| Field | Type | Notes |
|---|---|---|
model | string | Required. A catalog id such as deepseek/deepseek-v4.1-flash. List yours with GET /v1/models. |
messages | array | Required. Objects with role (system, user, assistant, tool) and content. |
temperature | number | Usually 0 to 2. Some reasoning models ignore or reject it. |
max_tokens | integer | Upper bound on generated tokens. Set it; see the note on reservations below. |
max_completion_tokens | integer | Newer OpenAI name for the same limit. Some models require it instead of max_tokens. |
stream | boolean | true returns Server-Sent Events. See streaming. |
stream_options.include_usage | boolean | With stream: true, adds a final chunk carrying usage. Off unless you ask. |
tools | array | Function definitions. See tool calling. |
tool_choice | string or object | "auto", "none", "required", or a specific function. Support varies by model. |
response_format | object | {"type": "json_object"} or a json_schema object, where the model supports it. |
n | integer | Number of choices, 1 to 4. Values outside that range return 400 invalid_request. |
stop, top_p, seed, presence_penalty, frequency_penalty | various | Passed through as sent. |
Parameters depend on the upstream model
Apart from model and n, the gateway does not validate or rewrite these fields. It passes them to the provider serving the model. If a model doesn't support response_format, tools or a given temperature, the provider decides what happens: it may ignore the field or reject the request. Check the model's page in the catalog before relying on a feature.
The request body can be up to 10 MB. Larger bodies return 413.
Example response#
{
"id": "chatcmpl-a1b2c3",
"object": "chat.completion",
"created": 1790000000,
"model": "deepseek/deepseek-v4.1-flash",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "429 Too Many Requests: the server is rate limiting you. Back off and retry after the Retry-After interval."
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 27,
"completion_tokens": 24,
"total_tokens": 51
}
}The body comes from the upstream provider, so the exact id format, the model string and any extra fields (such as reasoning_content or system_fingerprint) vary by model.
The usage object#
usage reports what the provider counted:
| Field | Meaning |
|---|---|
prompt_tokens | Input tokens, including any cached ones |
completion_tokens | Output tokens, including reasoning tokens on models that report them that way |
total_tokens | Sum of the two |
prompt_tokens_details.cached_tokens | Input tokens served from the provider's cache, when reported |
Billing uses these upstream counts, with cached input priced at the model's cache-read rate where one exists. If a provider sends no usage at all, the gateway estimates from the request and response size. Per-request costs appear in usage analytics; per-model rates are on the model catalog and pricing.
How max_tokens affects admission#
Before forwarding, the gateway reserves the worst-case cost of the request against your plan credits or wallet, using max_tokens (or max_completion_tokens) for the output side. If you leave it unset, the reservation assumes 8,192 output tokens. Two practical consequences:
- A key close to its monthly spend cap can be refused for a large
max_tokenswhile a small one still gets through. - If your balance can't cover the requested
max_tokens, the gateway may lower it to what the balance covers (never below 16). You'll seefinish_reason: "length"on a shorter answer. Top up in billing or lowermax_tokensyourself.
You are charged for actual usage, not the reservation.
Errors#
Gateway errors use the OpenAI error shape with a code you can branch on, and every response carries an x-tokens-request-id header. The full table is in errors, and per-minute and concurrency limits are in rate limits.