Two read-only endpoints let scripts and tools answer "which models can I call?" and "how much do I have left?" without a browser session. Both use the same API key as inference and neither is billed.
List models with GET /v1/models#
curl https://tokens.bd/v1/models \
-H "Authorization: Bearer $TOKENS_API_KEY"{
"object": "list",
"data": [
{
"id": "deepseek/deepseek-v4.1-flash",
"object": "model",
"created": 1788000000,
"owned_by": "tokens",
"permission": [],
"root": "deepseek/deepseek-v4.1-flash",
"parent": null
}
]
}| Field | Meaning |
|---|---|
id | The exact string to send as model. Always provider/model form. |
object | Always "model". |
created | Unix timestamp of when the model was added to the catalog. |
owned_by | Always "tokens", whichever lab built the model. |
root, parent, permission | Present for OpenAI SDK compatibility; root equals id. |
The response does not include prices, context windows or capabilities. Those are on the model catalog, one page per model.
Why a model is missing from the list#
The list is filtered for the key that calls it, so two keys on the same account can see different lists:
- Key allow-list. If the key was created with an allowed-models list, only those models appear.
- Plan and wallet. Each model is enabled for certain plan tiers and for pay-as-you-go. A model appears if your active plan's tier includes it, or if pay-as-you-go is allowed for it and your wallet balance is above zero.
- Catalog status. Only active catalog models are listed.
An empty data array usually means no active plan and no wallet balance. Subscribe or top up in billing; details in plans and wallet.
Check remaining usage with GET /v1/tokens/usage#
This endpoint is specific to Tokens. It returns the current plan, plan usage windows, wallet balance and the calling key's limits. The tokens.mjs usage CLI command reads it, and it's handy in a status bar or a pre-flight check in a long-running agent.
curl https://tokens.bd/v1/tokens/usage \
-H "Authorization: Bearer $TOKENS_API_KEY"{
"object": "tokens.usage",
"plan": {
"name": "Pro Monthly",
"tier": "monthly",
"periodEnd": "2026-11-02T08:15:00.000Z"
},
"windows": [
{
"type": "session_5h",
"label": "5-Hour Session",
"unit": "usd",
"limit": 5,
"used": 1.284,
"remaining": 3.716,
"percentUsed": 26,
"resetsAt": "2026-10-03T14:40:00.000Z"
}
],
"wallet": { "balanceUsd": 12.5 },
"key": {
"monthlySpendCapUsd": 20,
"allowedModels": null
}
}The plan name and numbers above are illustrative; yours depend on your plan.
| Field | Type | Meaning |
|---|---|---|
plan | object or null | Active subscription, or null on pay-as-you-go only |
plan.tier | string | Plan tier, such as weekly or monthly |
plan.periodEnd | ISO 8601 or null | When the current subscription period ends |
windows[] | array | Plan usage windows that have usage in their current period |
windows[].type | string | session_5h (rolling 5 hours), weekly or monthly |
windows[].unit | string | usd (credits, expressed in USD) or requests |
windows[].limit, used, remaining | number | In the window's unit |
windows[].percentUsed | integer | 0 to 100 |
windows[].resetsAt | ISO 8601 | When the window resets |
wallet | object or null | balanceUsd: prepaid wallet balance in USD, or null if there is no wallet |
key.monthlySpendCapUsd | number or null | The calling key's monthly cap, null if uncapped |
key.allowedModels | array or null | The key's allow-list, null if any model is allowed |
Note
A window that hasn't been used yet in its current period is left out of windows, so an empty array on a fresh plan or right after a reset is normal.
When a window is used up, inference requests return 429 window_exhausted until resetsAt. The dashboard can also notify you at 50, 75, 90 and 100 percent of a window; see usage and alerts.
The response is sent with Cache-Control: no-store. Polling it once a minute is plenty; it counts toward neither your bill nor your per-minute request limit.
Embeddings: POST /v1/embeddings#
The route exists and follows OpenAI's embeddings format, but it only works for catalog models that are embedding models. Chat models will fail on it. Check the model catalog for an embedding model before building on this endpoint; if none is listed, there is no embedding model available for your account yet.
Legacy completions: POST /v1/completions#
The old prompt-in, text-out completions endpoint is passed through for tools that still use it. It only works when the upstream behind a model supports it, and many chat models don't. Use chat completions for anything new.
Endpoints that don't exist#
Any other path under /v1 returns 404 with code unsupported_endpoint. That includes images, audio, files, batches, assistants, fine-tuning and moderations. See errors for the full code list.