# Request ids and debugging

> The request id headers on every Tokens response, how to quote one to support, what to log, how to reproduce a failing call with curl, and how to read the usage page and the error body.

When a request fails, or costs something you did not expect, you need to point at that one call. On Tokens, every response carries a request id, and every request that was billed gets a row on the usage page with the same id. This page shows where the id is, what to log around it, how to turn a failure from your app into a curl command anyone can run, and how to read the two places that tell you what happened: the error body and the usage page.

## The request id headers

Every response from `/v1`, success or error, carries these headers:

| Header                | Value                                                                                           |
| --------------------- | ----------------------------------------------------------------------------------------------- |
| `x-tokens-request-id` | The gateway's id for this request (a UUID). Always generated by Tokens. This is the one to quote. |
| `x-request-id`        | The `x-request-id` you sent, unchanged, or the same id as above if you sent none                |
| `x-trace-id`          | On successful inference responses: the same value as `x-tokens-request-id`                      |

Error bodies repeat the id as `request_id`. On `/v1/chat/completions`, `/v1/responses` and the other OpenAI-format endpoints it is inside `error`. On `/v1/messages`, which uses Anthropic's error format, it is at the top level of the body, next to `error`:

```json
{
  "error": {
    "message": "Prepaid wallet balance is insufficient for this request.",
    "type": "insufficient_quota",
    "code": "insufficient_credits",
    "param": null,
    "request_id": "8f0c7a4e-2b1d-4c55-9a51-3f7e2d9b6c10"
  }
}
```

```json
{
  "type": "error",
  "error": {
    "type": "billing_error",
    "message": "Prepaid wallet balance is insufficient for this request.",
    "code": "insufficient_credits"
  },
  "request_id": "8f0c7a4e-2b1d-4c55-9a51-3f7e2d9b6c10"
}
```

Some things to know:

- **Tokens does not pass provider headers through.** Only the content type, `cache-control`, `retry-after`, `x-request-id` and `x-tokens-*` headers reach you. A request id from the model's own provider never appears, so quote the Tokens id.
- **The id arrives with the headers.** For a streamed answer it is available before the first token. Log it at the start of the call, so you have it if the stream dies later.
- **Your own `x-request-id` is echoed back as is.** Use it to match a line in your logs to Tokens' id. Send a fresh one for each attempt, not one per user action, so a retry and its first try can be told apart.
- **A timeout on your side leaves you with no Tokens id,** because no response arrived. That is why you also log your own id before sending.
- **SDK helpers can show your id, not ours.** Some SDKs expose a `request_id` read from `x-request-id`. If you send your own `x-request-id`, that value is yours. Read `x-tokens-request-id` from the raw headers to get the gateway's id.

## Read the id in your code

:::code-tabs

```bash title="cURL"
curl -sS -i https://tokens.bd/v1/chat/completions \
  -H "Authorization: Bearer $TOKENS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"deepseek/deepseek-v4.1-flash","messages":[{"role":"user","content":"ping"}],"max_tokens":20}' \
  | grep -i "^x-tokens-request-id"
```

```python title="Python"
import os
import openai
from openai import OpenAI

client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"])

try:
    raw = client.chat.completions.with_raw_response.create(
        model="deepseek/deepseek-v4.1-flash",
        messages=[{"role": "user", "content": "ping"}],
        max_tokens=20,
    )
    print("request id:", raw.headers.get("x-tokens-request-id"))
    completion = raw.parse()
except openai.APIStatusError as e:
    print("failed:", e.status_code, e.code, e.response.headers.get("x-tokens-request-id"))
    raise
```

```typescript title="Node.js"
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://tokens.bd/v1",
  apiKey: process.env.TOKENS_API_KEY,
});

try {
  const { data, response } = await client.chat.completions
    .create({
      model: "deepseek/deepseek-v4.1-flash",
      messages: [{ role: "user", content: "ping" }],
      max_tokens: 20,
    })
    .withResponse();
  console.log("request id:", response.headers.get("x-tokens-request-id"));
  console.log(data.choices[0]?.message.content);
} catch (err) {
  if (err instanceof OpenAI.APIError) {
    const headers = err.headers as Headers | Record<string, string | undefined> | undefined;
    const id = headers instanceof Headers ? headers.get("x-tokens-request-id") : headers?.["x-tokens-request-id"];
    console.error("failed:", err.status, err.code, id);
  }
  throw err;
}
```

:::

With the Anthropic SDKs the same headers are on the raw response; each SDK documents how to reach it. For a `fetch` or `requests` call, read the response headers directly.

## What to log

One structured line per attempt is enough. Log it when the response headers arrive, and again if something fails later.

| Field                       | Why                                                              |
| --------------------------- | ---------------------------------------------------------------- |
| `x-tokens-request-id`       | Lets support find the request                                    |
| Your own request id         | Ties the call to your own logs and to a retry sequence           |
| Time, with time zone (UTC)  | Fallback when an id is missing                                   |
| Endpoint and model          | `chat/completions`, `messages`, and the exact model id           |
| HTTP status and `error.code` | What happened. Branch on the code, not the message               |
| Latency, and time to first token for streams | Separates slow providers from slow clients          |
| `usage` from the response   | Tokens in, cached tokens, tokens out, so you can explain a cost   |
| Attempt number              | Shows retries                                                    |
| Which key (name, never the secret) | Finds the right key when you have several                |

Do not log:

- **The API key,** in any form. Not in headers, not in a dumped config.
- **Full prompts and answers by default.** They can contain customer data. Tokens itself stores usage metadata by request id, not prompt content, so the id is what finds a request. If you log bodies to debug, do it with a short retention and redaction.

## Quote an id to support

When you open a ticket in [support](/dashboard/support), include:

1. The `x-tokens-request-id`, or several if it is a pattern.
2. When it happened, with the time zone.
3. The endpoint and the model id.
4. The HTTP status and `error.code`, or, for a billing question, what you expected to be charged.
5. Whether it is reproducible, and the curl command if you have one (next section).

Never include the API key. The id lets support find usage metadata for your call. [Getting help](/docs/support) explains the rest of the process, and the [status page](/status) shows whether something is broken for everyone.

## Reproduce a failing request with curl

If a request fails in your app, take the app out of the picture. A curl command that fails the same way settles whether the problem is your code, your configuration, the key or the account.

1. **Capture the exact body** your app sent: the JSON for `model`, `messages`, `max_tokens`, tools and the rest. Remove customer text you do not want to share.
2. **Save it to a file** and send it with the same key and the same endpoint.
3. **Save headers and body separately** so the id and the error survive.

:::code-tabs

```bash title="Chat completions"
curl -sS -D headers.txt -o response.json -w "HTTP %{http_code}\n" \
  https://tokens.bd/v1/chat/completions \
  -H "Authorization: Bearer $TOKENS_API_KEY" \
  -H "Content-Type: application/json" \
  -H "x-request-id: repro-$(date +%s)" \
  -d @body.json

grep -i "^x-tokens-request-id" headers.txt
jq '.error // .' response.json
```

```bash title="Messages"
curl -sS -D headers.txt -o response.json -w "HTTP %{http_code}\n" \
  https://tokens.bd/v1/messages \
  -H "x-api-key: $TOKENS_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -H "x-request-id: repro-$(date +%s)" \
  -d @body.json

grep -i "^x-tokens-request-id" headers.txt
jq '.error // .' response.json
```

```powershell title="Windows PowerShell"
curl.exe -sS -D headers.txt -o response.json -w "HTTP %{http_code}`n" `
  https://tokens.bd/v1/chat/completions `
  -H "Authorization: Bearer $env:TOKENS_API_KEY" `
  -H "Content-Type: application/json" `
  -d "@body.json"

Select-String -Path headers.txt -Pattern "x-tokens-request-id"
Get-Content response.json
```

:::

In Windows PowerShell 5.1 `curl` is an alias for another command, so use `curl.exe`. To watch a streamed answer arrive, add `-N` and `"stream": true` to the body.

Then narrow it down, changing one thing at a time:

| If the curl call...               | It points to                                                                                 |
| --------------------------------- | -------------------------------------------------------------------------------------------- |
| Fails the same way                | The request itself, the key or the account. Read `error.code` and look it up in [errors](/docs/errors) |
| Works                             | Your app: a different key in the environment, a changed body, a proxy, a timeout, or a library default |
| Works without `stream`, fails with it | A buffering proxy, or a client that does not handle SSE. See [streaming](/docs/streaming) |
| Fails only with a certain model   | That model's parameters or availability. Try another from `GET /v1/models`                   |
| Fails only with tools or a large body | The tool schema, the 10 MB body limit, or the model's context window                     |
| Fails only some of the time       | A rate limit, or a provider problem. Check `Retry-After` and the [status page](/status)      |

A quick sanity check that the key and the base URL are right is `GET /v1/models`. It does not run a model and does not count toward your per-minute limit.

## Read the error body

Branch on `error.code`. Messages are for people and can change. The shape and the codes are in [errors](/docs/errors). A short way to read one:

| Part                | Read it as                                                                                     |
| ------------------- | ---------------------------------------------------------------------------------------------- |
| HTTP status         | The class: 4xx is about the request, key or account, 5xx is the gateway or a provider         |
| `error.code`        | The exact reason. `insufficient_credits`, `window_exhausted`, `model_not_found` and so on      |
| `error.message`     | Context: for example the list of models a key may use, or when a window resets                 |
| `request_id`        | What you quote to support                                                                      |
| `Retry-After` header | How long to wait, in seconds, on a 429                                                        |

Two behaviors that confuse people:

- **Provider errors are generic.** When the model's provider rejects or fails a request, the message is replaced with a generic one so provider internals do not leak. The status and `code` still say what kind of failure it was. For a 400 from a provider, look at your parameters and at whether the model supports them.
- **A stream can fail after a 200.** Once streaming has started the status is already 200. If the provider fails midway, the connection closes without `data: [DONE]` (or without `message_stop` on Messages), and there is no error body. Treat a stream with no finish reason as incomplete, and use the request id you logged at the start.

## Read the usage page

[Usage](/dashboard/usage) lists your recent requests, newest first, in an activity table:

| Column   | What it shows                                                                                              |
| -------- | ---------------------------------------------------------------------------------------------------------- |
| Time     | When the request was recorded                                                                              |
| Usage    | The cost of the request                                                                                    |
| Tokens   | `input · N cached · output`. The cached part, shown only when there is one, is cache reads and writes together |
| Timing   | How long the request took                                                                                  |
| Model    | The model id you called                                                                                    |
| Mode     | `api` for requests through `/v1`                                                                           |
| Status   | `COMPLETED` for a billed request                                                                           |
| Trace ID | The request id. Click it to copy the full value; the table shows only the first eight characters           |

The Trace ID of an API request is the same value as its `x-tokens-request-id`. To find one request, copy the id from your logs and look for it in this column. The table has pages, so for an older request use the per-page selector at the bottom to show more rows.

What this page can and cannot tell you:

- **Only billed requests appear.** The gateway records usage when a request has run and been billed. A request that was rejected before running (bad key, no balance, a rate limit) or that failed at the provider has no row. For those, your own log of the request id and the error body is the record.
- **A request you cancelled still appears.** If your client disconnected mid-stream, you are billed for the input and the output generated so far, and the row shows those numbers.
- **An Estimated badge means no usage was reported.** If the provider sent no token counts, Tokens estimates them from the size of the request and the answer and marks the row. The estimate is what you were billed.
- **Cached tokens explain a low cost.** If a long prompt cost little, check the cached part of the Tokens column. See [prompt caching](/docs/prompt-caching).
- **A big cost from a small prompt has usual causes:** a long conversation history sent again on every turn, a large `max_tokens` on a model that uses it, reasoning tokens counted as output, retries that were billed more than once, or several requests from a loop. The per-request tokens tell you which.

The totals, the daily chart and the plan windows are explained in [usage and alerts](/docs/usage-and-alerts). For a check from code, `GET /v1/tokens/usage` returns your windows, balance and the key's cap without being billed.

## A short debugging routine

1. Read `error.code` and the status. Look the code up in [errors](/docs/errors).
2. Get the `x-tokens-request-id` from the response, or from your log.
3. Reproduce with curl and a saved body. Change one thing at a time.
4. Check [status](/status) if several requests or models fail at once.
5. Look at [Usage](/dashboard/usage) if the question is about cost or tokens.
6. If it is still unexplained, open a ticket with the id, the time, the model, the status and the code.

---
Page: https://tokens.bd/docs/request-ids-and-debugging
