Skip to content

Chat Completions

POST /v1/chat/completions: request fields, a full request and response, the usage object, and how the gateway treats max_tokens, n and streaming.

On this page

POST https://tokens.bd/v1/chat/completions is the OpenAI-compatible chat completions API, and the endpoint most SDKs and coding agents use. The request and response follow OpenAI's format; the gateway checks your key, plan and limits, then forwards the body to the upstream provider for the model you named.

Make a chat completions request#

curl https://tokens.bd/v1/chat/completions \
  -H "Authorization: Bearer $TOKENS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek/deepseek-v4.1-flash",
    "messages": [
      {"role": "system", "content": "You are a concise senior engineer."},
      {"role": "user", "content": "What does HTTP 429 mean?"}
    ],
    "temperature": 0.2,
    "max_tokens": 300
  }'

The Content-Type: application/json header matters. Without it the gateway can't read the body, and you get a 400 saying the request must specify a model field even though it does.

Request fields#

FieldTypeNotes
modelstringRequired. A catalog id such as deepseek/deepseek-v4.1-flash. List yours with GET /v1/models.
messagesarrayRequired. Objects with role (system, user, assistant, tool) and content.
temperaturenumberUsually 0 to 2. Some reasoning models ignore or reject it.
max_tokensintegerUpper bound on generated tokens. Set it; see the note on reservations below.
max_completion_tokensintegerNewer OpenAI name for the same limit. Some models require it instead of max_tokens.
streambooleantrue returns Server-Sent Events. See streaming.
stream_options.include_usagebooleanWith stream: true, adds a final chunk carrying usage. Off unless you ask.
toolsarrayFunction definitions. See tool calling.
tool_choicestring or object"auto", "none", "required", or a specific function. Support varies by model.
response_formatobject{"type": "json_object"} or a json_schema object, where the model supports it.
nintegerNumber of choices, 1 to 4. Values outside that range return 400 invalid_request.
stop, top_p, seed, presence_penalty, frequency_penaltyvariousPassed through as sent.

Parameters depend on the upstream model

Apart from model and n, the gateway does not validate or rewrite these fields. It passes them to the provider serving the model. If a model doesn't support response_format, tools or a given temperature, the provider decides what happens: it may ignore the field or reject the request. Check the model's page in the catalog before relying on a feature.

The request body can be up to 10 MB. Larger bodies return 413.

Example response#

json
{
  "id": "chatcmpl-a1b2c3",
  "object": "chat.completion",
  "created": 1790000000,
  "model": "deepseek/deepseek-v4.1-flash",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "429 Too Many Requests: the server is rate limiting you. Back off and retry after the Retry-After interval."
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 27,
    "completion_tokens": 24,
    "total_tokens": 51
  }
}

The body comes from the upstream provider, so the exact id format, the model string and any extra fields (such as reasoning_content or system_fingerprint) vary by model.

The usage object#

usage reports what the provider counted:

FieldMeaning
prompt_tokensInput tokens, including any cached ones
completion_tokensOutput tokens, including reasoning tokens on models that report them that way
total_tokensSum of the two
prompt_tokens_details.cached_tokensInput tokens served from the provider's cache, when reported

Billing uses these upstream counts, with cached input priced at the model's cache-read rate where one exists. If a provider sends no usage at all, the gateway estimates from the request and response size. Per-request costs appear in usage analytics; per-model rates are on the model catalog and pricing.

How max_tokens affects admission#

Before forwarding, the gateway reserves the worst-case cost of the request against your plan credits or wallet, using max_tokens (or max_completion_tokens) for the output side. If you leave it unset, the reservation assumes 8,192 output tokens. Two practical consequences:

  • A key close to its monthly spend cap can be refused for a large max_tokens while a small one still gets through.
  • If your balance can't cover the requested max_tokens, the gateway may lower it to what the balance covers (never below 16). You'll see finish_reason: "length" on a shorter answer. Top up in billing or lower max_tokens yourself.

You are charged for actual usage, not the reservation.

Errors#

Gateway errors use the OpenAI error shape with a code you can branch on, and every response carries an x-tokens-request-id header. The full table is in errors, and per-minute and concurrency limits are in rate limits.

Was this page helpful?

Still stuck? Open a support ticket

Need help configuring your agent?

Test your connection with the connection tester, or create an API key.