Skip to content
Start here

Platform overview

What Tokens is, how a request travels from your tool to the model provider, what gets metered, and the handful of concepts you need before your first call.

On this page

Tokens is an AI model gateway: one API key gives you OpenAI-compatible and Anthropic-compatible endpoints for models from many providers, and you pay in USD or BDT. This page explains how the platform works so the rest of the docs make sense.

What Tokens is#

Most coding agents and SDKs already speak one of two dialects: the OpenAI API or the Anthropic Messages API. Tokens speaks both. You point your tool at Tokens instead of at a single provider, use one tok_live_ key, and pick any model your plan or wallet covers by changing the model field.

You useBase URLTypical clients
OpenAI-compatible APIhttps://tokens.bd/v1OpenAI SDKs, OpenCode, Codex CLI, Cursor, Cline, Aider, Continue
Anthropic-compatible APIhttps://tokens.bdClaude Code, the Anthropic SDK (it appends /v1/messages itself)

The full model list with prices is the model catalog. Plans and top-ups are on pricing.

How a request flows#

  1. Your tool sends a request to https://tokens.bd/v1/... with your key in Authorization: Bearer or x-api-key.
  2. Tokens checks it: the key is valid and active, the model is allowed on that key and on your plan, you are inside your rate limits and usage windows, and you have credit to pay for it.
  3. Tokens forwards it to an upstream source for that model. If that source answers with 429, 502, 503, 504 or drops the connection, Tokens retries on another source for the same model automatically.
  4. The response streams back to your tool unchanged (SSE for streaming requests), and Tokens meters the tokens as they pass.
  5. The cost is settled against your plan credits or wallet when the response finishes.

Every response carries an x-tokens-request-id header. Keep it when something goes wrong; it is the fastest way for support to find your request.

Note

Tokens adds one network hop between you and the provider. Long requests are fine (response headers can take up to 600 seconds), but don't expect lower latency than calling the provider directly.

What the API covers#

Supported endpoints: POST /v1/chat/completions, POST /v1/messages, POST /v1/responses, POST /v1/completions (legacy), POST /v1/embeddings (only for models that are embedding models), GET /v1/models and GET /v1/tokens/usage.

Not supported: image generation, audio, files, batches, assistants, fine-tuning and moderations. Calls to those return 404 unsupported_endpoint. Browser calls are not supported either, because responses carry no CORS headers. Call Tokens from a server, a script or a CLI tool.

What is metered#

Tokens meters every inference request, from /v1 and from the dashboard playground. For each request it records:

  • the model, input tokens, output tokens and cache-read tokens
  • the cost at that model's catalog price
  • latency, timestamps, the key used and the request id

That metadata is what you see in the usage dashboard and the CSV export. Prompt and response content is not stored. See security and data privacy for the details.

Read-only calls such as GET /v1/tokens/usage are not metered.

Key concepts#

API key#

A secret that starts with tok_live_. It is shown once when you create it in /dashboard/keys. A key can carry an optional monthly spend cap and an optional list of allowed models. See API keys.

Model id#

Models are addressed by an alias in provider/model form, for example deepseek/deepseek-v4.1-flash. The alias is what goes in the model field. It does not always match the provider's own id, so copy it from /models or from GET /v1/models, which returns exactly the models your key can use. Choosing a model helps you pick one.

Plan vs wallet#

There are two ways to pay for usage, and you can have both:

  • A plan is a weekly or monthly subscription that gives you a credit allowance (100 credits = 1 USD by default) and may include usage windows.
  • The wallet is a prepaid USD balance for pay-as-you-go use. You can top it up in USD or BDT.

Details are in plans, credits and wallet.

Usage windows#

Some plans spread their allowance over time with windows: a rolling 5-hour session, a weekly ceiling and a monthly ceiling. When a window is used up, requests return 429 window_exhausted with a Retry-After header that says how long until it resets. You can see every window in the dashboard, with GET /v1/tokens/usage, or with the CLI's usage command. See usage, limits and alerts.

Where to go next#

Was this page helpful?

Still stuck? Open a support ticket

Need help configuring your agent?

Test your connection with the connection tester, or create an API key.