Skip to content

Request ids and debugging

The request id headers on every Tokens response, how to quote one to support, what to log, how to reproduce a failing call with curl, and how to read the usage page and the error body.

On this page

When a request fails, or costs something you did not expect, you need to point at that one call. On Tokens, every response carries a request id, and every request that was billed gets a row on the usage page with the same id. This page shows where the id is, what to log around it, how to turn a failure from your app into a curl command anyone can run, and how to read the two places that tell you what happened: the error body and the usage page.

The request id headers#

Every response from /v1, success or error, carries these headers:

HeaderValue
x-tokens-request-idThe gateway's id for this request (a UUID). Always generated by Tokens. This is the one to quote.
x-request-idThe x-request-id you sent, unchanged, or the same id as above if you sent none
x-trace-idOn successful inference responses: the same value as x-tokens-request-id

Error bodies repeat the id as request_id. On /v1/chat/completions, /v1/responses and the other OpenAI-format endpoints it is inside error. On /v1/messages, which uses Anthropic's error format, it is at the top level of the body, next to error:

json
{
  "error": {
    "message": "Prepaid wallet balance is insufficient for this request.",
    "type": "insufficient_quota",
    "code": "insufficient_credits",
    "param": null,
    "request_id": "8f0c7a4e-2b1d-4c55-9a51-3f7e2d9b6c10"
  }
}
json
{
  "type": "error",
  "error": {
    "type": "billing_error",
    "message": "Prepaid wallet balance is insufficient for this request.",
    "code": "insufficient_credits"
  },
  "request_id": "8f0c7a4e-2b1d-4c55-9a51-3f7e2d9b6c10"
}

Some things to know:

  • Tokens does not pass provider headers through. Only the content type, cache-control, retry-after, x-request-id and x-tokens-* headers reach you. A request id from the model's own provider never appears, so quote the Tokens id.
  • The id arrives with the headers. For a streamed answer it is available before the first token. Log it at the start of the call, so you have it if the stream dies later.
  • Your own x-request-id is echoed back as is. Use it to match a line in your logs to Tokens' id. Send a fresh one for each attempt, not one per user action, so a retry and its first try can be told apart.
  • A timeout on your side leaves you with no Tokens id, because no response arrived. That is why you also log your own id before sending.
  • SDK helpers can show your id, not ours. Some SDKs expose a request_id read from x-request-id. If you send your own x-request-id, that value is yours. Read x-tokens-request-id from the raw headers to get the gateway's id.

Read the id in your code#

curl -sS -i https://tokens.bd/v1/chat/completions \
  -H "Authorization: Bearer $TOKENS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"deepseek/deepseek-v4.1-flash","messages":[{"role":"user","content":"ping"}],"max_tokens":20}' \
  | grep -i "^x-tokens-request-id"

With the Anthropic SDKs the same headers are on the raw response; each SDK documents how to reach it. For a fetch or requests call, read the response headers directly.

What to log#

One structured line per attempt is enough. Log it when the response headers arrive, and again if something fails later.

FieldWhy
x-tokens-request-idLets support find the request
Your own request idTies the call to your own logs and to a retry sequence
Time, with time zone (UTC)Fallback when an id is missing
Endpoint and modelchat/completions, messages, and the exact model id
HTTP status and error.codeWhat happened. Branch on the code, not the message
Latency, and time to first token for streamsSeparates slow providers from slow clients
usage from the responseTokens in, cached tokens, tokens out, so you can explain a cost
Attempt numberShows retries
Which key (name, never the secret)Finds the right key when you have several

Do not log:

  • The API key, in any form. Not in headers, not in a dumped config.
  • Full prompts and answers by default. They can contain customer data. Tokens itself stores usage metadata by request id, not prompt content, so the id is what finds a request. If you log bodies to debug, do it with a short retention and redaction.

Quote an id to support#

When you open a ticket in support, include:

  1. The x-tokens-request-id, or several if it is a pattern.
  2. When it happened, with the time zone.
  3. The endpoint and the model id.
  4. The HTTP status and error.code, or, for a billing question, what you expected to be charged.
  5. Whether it is reproducible, and the curl command if you have one (next section).

Never include the API key. The id lets support find usage metadata for your call. Getting help explains the rest of the process, and the status page shows whether something is broken for everyone.

Reproduce a failing request with curl#

If a request fails in your app, take the app out of the picture. A curl command that fails the same way settles whether the problem is your code, your configuration, the key or the account.

  1. Capture the exact body your app sent: the JSON for model, messages, max_tokens, tools and the rest. Remove customer text you do not want to share.
  2. Save it to a file and send it with the same key and the same endpoint.
  3. Save headers and body separately so the id and the error survive.
curl -sS -D headers.txt -o response.json -w "HTTP %{http_code}\n" \
  https://tokens.bd/v1/chat/completions \
  -H "Authorization: Bearer $TOKENS_API_KEY" \
  -H "Content-Type: application/json" \
  -H "x-request-id: repro-$(date +%s)" \
  -d @body.json

grep -i "^x-tokens-request-id" headers.txt
jq '.error // .' response.json

In Windows PowerShell 5.1 curl is an alias for another command, so use curl.exe. To watch a streamed answer arrive, add -N and "stream": true to the body.

Then narrow it down, changing one thing at a time:

If the curl call...It points to
Fails the same wayThe request itself, the key or the account. Read error.code and look it up in errors
WorksYour app: a different key in the environment, a changed body, a proxy, a timeout, or a library default
Works without stream, fails with itA buffering proxy, or a client that does not handle SSE. See streaming
Fails only with a certain modelThat model's parameters or availability. Try another from GET /v1/models
Fails only with tools or a large bodyThe tool schema, the 10 MB body limit, or the model's context window
Fails only some of the timeA rate limit, or a provider problem. Check Retry-After and the status page

A quick sanity check that the key and the base URL are right is GET /v1/models. It does not run a model and does not count toward your per-minute limit.

Read the error body#

Branch on error.code. Messages are for people and can change. The shape and the codes are in errors. A short way to read one:

PartRead it as
HTTP statusThe class: 4xx is about the request, key or account, 5xx is the gateway or a provider
error.codeThe exact reason. insufficient_credits, window_exhausted, model_not_found and so on
error.messageContext: for example the list of models a key may use, or when a window resets
request_idWhat you quote to support
Retry-After headerHow long to wait, in seconds, on a 429

Two behaviors that confuse people:

  • Provider errors are generic. When the model's provider rejects or fails a request, the message is replaced with a generic one so provider internals do not leak. The status and code still say what kind of failure it was. For a 400 from a provider, look at your parameters and at whether the model supports them.
  • A stream can fail after a 200. Once streaming has started the status is already 200. If the provider fails midway, the connection closes without data: [DONE] (or without message_stop on Messages), and there is no error body. Treat a stream with no finish reason as incomplete, and use the request id you logged at the start.

Read the usage page#

Usage lists your recent requests, newest first, in an activity table:

ColumnWhat it shows
TimeWhen the request was recorded
UsageThe cost of the request
Tokensinput · N cached · output. The cached part, shown only when there is one, is cache reads and writes together
TimingHow long the request took
ModelThe model id you called
Modeapi for requests through /v1
StatusCOMPLETED for a billed request
Trace IDThe request id. Click it to copy the full value; the table shows only the first eight characters

The Trace ID of an API request is the same value as its x-tokens-request-id. To find one request, copy the id from your logs and look for it in this column. The table has pages, so for an older request use the per-page selector at the bottom to show more rows.

What this page can and cannot tell you:

  • Only billed requests appear. The gateway records usage when a request has run and been billed. A request that was rejected before running (bad key, no balance, a rate limit) or that failed at the provider has no row. For those, your own log of the request id and the error body is the record.
  • A request you cancelled still appears. If your client disconnected mid-stream, you are billed for the input and the output generated so far, and the row shows those numbers.
  • An Estimated badge means no usage was reported. If the provider sent no token counts, Tokens estimates them from the size of the request and the answer and marks the row. The estimate is what you were billed.
  • Cached tokens explain a low cost. If a long prompt cost little, check the cached part of the Tokens column. See prompt caching.
  • A big cost from a small prompt has usual causes: a long conversation history sent again on every turn, a large max_tokens on a model that uses it, reasoning tokens counted as output, retries that were billed more than once, or several requests from a loop. The per-request tokens tell you which.

The totals, the daily chart and the plan windows are explained in usage and alerts. For a check from code, GET /v1/tokens/usage returns your windows, balance and the key's cap without being billed.

A short debugging routine#

  1. Read error.code and the status. Look the code up in errors.
  2. Get the x-tokens-request-id from the response, or from your log.
  3. Reproduce with curl and a saved body. Change one thing at a time.
  4. Check status if several requests or models fail at once.
  5. Look at Usage if the question is about cost or tokens.
  6. If it is still unexplained, open a ticket with the id, the time, the model, the status and the code.

Was this page helpful?

Still stuck? Open a support ticket

Need help configuring your agent?

Test your connection with the connection tester, or create an API key.