When a request fails, or costs something you did not expect, you need to point at that one call. On Tokens, every response carries a request id, and every request that was billed gets a row on the usage page with the same id. This page shows where the id is, what to log around it, how to turn a failure from your app into a curl command anyone can run, and how to read the two places that tell you what happened: the error body and the usage page.
The request id headers#
Every response from /v1, success or error, carries these headers:
| Header | Value |
|---|---|
x-tokens-request-id | The gateway's id for this request (a UUID). Always generated by Tokens. This is the one to quote. |
x-request-id | The x-request-id you sent, unchanged, or the same id as above if you sent none |
x-trace-id | On successful inference responses: the same value as x-tokens-request-id |
Error bodies repeat the id as request_id. On /v1/chat/completions, /v1/responses and the other OpenAI-format endpoints it is inside error. On /v1/messages, which uses Anthropic's error format, it is at the top level of the body, next to error:
{
"error": {
"message": "Prepaid wallet balance is insufficient for this request.",
"type": "insufficient_quota",
"code": "insufficient_credits",
"param": null,
"request_id": "8f0c7a4e-2b1d-4c55-9a51-3f7e2d9b6c10"
}
}{
"type": "error",
"error": {
"type": "billing_error",
"message": "Prepaid wallet balance is insufficient for this request.",
"code": "insufficient_credits"
},
"request_id": "8f0c7a4e-2b1d-4c55-9a51-3f7e2d9b6c10"
}Some things to know:
- Tokens does not pass provider headers through. Only the content type,
cache-control,retry-after,x-request-idandx-tokens-*headers reach you. A request id from the model's own provider never appears, so quote the Tokens id. - The id arrives with the headers. For a streamed answer it is available before the first token. Log it at the start of the call, so you have it if the stream dies later.
- Your own
x-request-idis echoed back as is. Use it to match a line in your logs to Tokens' id. Send a fresh one for each attempt, not one per user action, so a retry and its first try can be told apart. - A timeout on your side leaves you with no Tokens id, because no response arrived. That is why you also log your own id before sending.
- SDK helpers can show your id, not ours. Some SDKs expose a
request_idread fromx-request-id. If you send your ownx-request-id, that value is yours. Readx-tokens-request-idfrom the raw headers to get the gateway's id.
Read the id in your code#
curl -sS -i https://tokens.bd/v1/chat/completions \
-H "Authorization: Bearer $TOKENS_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"deepseek/deepseek-v4.1-flash","messages":[{"role":"user","content":"ping"}],"max_tokens":20}' \
| grep -i "^x-tokens-request-id"import os
import openai
from openai import OpenAI
client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"])
try:
raw = client.chat.completions.with_raw_response.create(
model="deepseek/deepseek-v4.1-flash",
messages=[{"role": "user", "content": "ping"}],
max_tokens=20,
)
print("request id:", raw.headers.get("x-tokens-request-id"))
completion = raw.parse()
except openai.APIStatusError as e:
print("failed:", e.status_code, e.code, e.response.headers.get("x-tokens-request-id"))
raiseimport OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://tokens.bd/v1",
apiKey: process.env.TOKENS_API_KEY,
});
try {
const { data, response } = await client.chat.completions
.create({
model: "deepseek/deepseek-v4.1-flash",
messages: [{ role: "user", content: "ping" }],
max_tokens: 20,
})
.withResponse();
console.log("request id:", response.headers.get("x-tokens-request-id"));
console.log(data.choices[0]?.message.content);
} catch (err) {
if (err instanceof OpenAI.APIError) {
const headers = err.headers as Headers | Record<string, string | undefined> | undefined;
const id = headers instanceof Headers ? headers.get("x-tokens-request-id") : headers?.["x-tokens-request-id"];
console.error("failed:", err.status, err.code, id);
}
throw err;
}With the Anthropic SDKs the same headers are on the raw response; each SDK documents how to reach it. For a fetch or requests call, read the response headers directly.
What to log#
One structured line per attempt is enough. Log it when the response headers arrive, and again if something fails later.
| Field | Why |
|---|---|
x-tokens-request-id | Lets support find the request |
| Your own request id | Ties the call to your own logs and to a retry sequence |
| Time, with time zone (UTC) | Fallback when an id is missing |
| Endpoint and model | chat/completions, messages, and the exact model id |
HTTP status and error.code | What happened. Branch on the code, not the message |
| Latency, and time to first token for streams | Separates slow providers from slow clients |
usage from the response | Tokens in, cached tokens, tokens out, so you can explain a cost |
| Attempt number | Shows retries |
| Which key (name, never the secret) | Finds the right key when you have several |
Do not log:
- The API key, in any form. Not in headers, not in a dumped config.
- Full prompts and answers by default. They can contain customer data. Tokens itself stores usage metadata by request id, not prompt content, so the id is what finds a request. If you log bodies to debug, do it with a short retention and redaction.
Quote an id to support#
When you open a ticket in support, include:
- The
x-tokens-request-id, or several if it is a pattern. - When it happened, with the time zone.
- The endpoint and the model id.
- The HTTP status and
error.code, or, for a billing question, what you expected to be charged. - Whether it is reproducible, and the curl command if you have one (next section).
Never include the API key. The id lets support find usage metadata for your call. Getting help explains the rest of the process, and the status page shows whether something is broken for everyone.
Reproduce a failing request with curl#
If a request fails in your app, take the app out of the picture. A curl command that fails the same way settles whether the problem is your code, your configuration, the key or the account.
- Capture the exact body your app sent: the JSON for
model,messages,max_tokens, tools and the rest. Remove customer text you do not want to share. - Save it to a file and send it with the same key and the same endpoint.
- Save headers and body separately so the id and the error survive.
curl -sS -D headers.txt -o response.json -w "HTTP %{http_code}\n" \
https://tokens.bd/v1/chat/completions \
-H "Authorization: Bearer $TOKENS_API_KEY" \
-H "Content-Type: application/json" \
-H "x-request-id: repro-$(date +%s)" \
-d @body.json
grep -i "^x-tokens-request-id" headers.txt
jq '.error // .' response.jsoncurl -sS -D headers.txt -o response.json -w "HTTP %{http_code}\n" \
https://tokens.bd/v1/messages \
-H "x-api-key: $TOKENS_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-H "x-request-id: repro-$(date +%s)" \
-d @body.json
grep -i "^x-tokens-request-id" headers.txt
jq '.error // .' response.jsoncurl.exe -sS -D headers.txt -o response.json -w "HTTP %{http_code}`n" `
https://tokens.bd/v1/chat/completions `
-H "Authorization: Bearer $env:TOKENS_API_KEY" `
-H "Content-Type: application/json" `
-d "@body.json"
Select-String -Path headers.txt -Pattern "x-tokens-request-id"
Get-Content response.jsonIn Windows PowerShell 5.1 curl is an alias for another command, so use curl.exe. To watch a streamed answer arrive, add -N and "stream": true to the body.
Then narrow it down, changing one thing at a time:
| If the curl call... | It points to |
|---|---|
| Fails the same way | The request itself, the key or the account. Read error.code and look it up in errors |
| Works | Your app: a different key in the environment, a changed body, a proxy, a timeout, or a library default |
Works without stream, fails with it | A buffering proxy, or a client that does not handle SSE. See streaming |
| Fails only with a certain model | That model's parameters or availability. Try another from GET /v1/models |
| Fails only with tools or a large body | The tool schema, the 10 MB body limit, or the model's context window |
| Fails only some of the time | A rate limit, or a provider problem. Check Retry-After and the status page |
A quick sanity check that the key and the base URL are right is GET /v1/models. It does not run a model and does not count toward your per-minute limit.
Read the error body#
Branch on error.code. Messages are for people and can change. The shape and the codes are in errors. A short way to read one:
| Part | Read it as |
|---|---|
| HTTP status | The class: 4xx is about the request, key or account, 5xx is the gateway or a provider |
error.code | The exact reason. insufficient_credits, window_exhausted, model_not_found and so on |
error.message | Context: for example the list of models a key may use, or when a window resets |
request_id | What you quote to support |
Retry-After header | How long to wait, in seconds, on a 429 |
Two behaviors that confuse people:
- Provider errors are generic. When the model's provider rejects or fails a request, the message is replaced with a generic one so provider internals do not leak. The status and
codestill say what kind of failure it was. For a 400 from a provider, look at your parameters and at whether the model supports them. - A stream can fail after a 200. Once streaming has started the status is already 200. If the provider fails midway, the connection closes without
data: [DONE](or withoutmessage_stopon Messages), and there is no error body. Treat a stream with no finish reason as incomplete, and use the request id you logged at the start.
Read the usage page#
Usage lists your recent requests, newest first, in an activity table:
| Column | What it shows |
|---|---|
| Time | When the request was recorded |
| Usage | The cost of the request |
| Tokens | input · N cached · output. The cached part, shown only when there is one, is cache reads and writes together |
| Timing | How long the request took |
| Model | The model id you called |
| Mode | api for requests through /v1 |
| Status | COMPLETED for a billed request |
| Trace ID | The request id. Click it to copy the full value; the table shows only the first eight characters |
The Trace ID of an API request is the same value as its x-tokens-request-id. To find one request, copy the id from your logs and look for it in this column. The table has pages, so for an older request use the per-page selector at the bottom to show more rows.
What this page can and cannot tell you:
- Only billed requests appear. The gateway records usage when a request has run and been billed. A request that was rejected before running (bad key, no balance, a rate limit) or that failed at the provider has no row. For those, your own log of the request id and the error body is the record.
- A request you cancelled still appears. If your client disconnected mid-stream, you are billed for the input and the output generated so far, and the row shows those numbers.
- An Estimated badge means no usage was reported. If the provider sent no token counts, Tokens estimates them from the size of the request and the answer and marks the row. The estimate is what you were billed.
- Cached tokens explain a low cost. If a long prompt cost little, check the cached part of the Tokens column. See prompt caching.
- A big cost from a small prompt has usual causes: a long conversation history sent again on every turn, a large
max_tokenson a model that uses it, reasoning tokens counted as output, retries that were billed more than once, or several requests from a loop. The per-request tokens tell you which.
The totals, the daily chart and the plan windows are explained in usage and alerts. For a check from code, GET /v1/tokens/usage returns your windows, balance and the key's cap without being billed.
A short debugging routine#
- Read
error.codeand the status. Look the code up in errors. - Get the
x-tokens-request-idfrom the response, or from your log. - Reproduce with curl and a saved body. Change one thing at a time.
- Check status if several requests or models fail at once.
- Look at Usage if the question is about cost or tokens.
- If it is still unexplained, open a ticket with the id, the time, the model, the status and the code.