Skip to content

Models and Usage Endpoints

GET /v1/models lists the models your key can call. GET /v1/tokens/usage returns your plan, usage windows, wallet balance and key limits. Plus notes on embeddings and legacy completions.

On this page

Two read-only endpoints let scripts and tools answer "which models can I call?" and "how much do I have left?" without a browser session. Both use the same API key as inference and neither is billed.

List models with GET /v1/models#

bash
curl https://tokens.bd/v1/models \
  -H "Authorization: Bearer $TOKENS_API_KEY"
json
{
  "object": "list",
  "data": [
    {
      "id": "deepseek/deepseek-v4.1-flash",
      "object": "model",
      "created": 1788000000,
      "owned_by": "tokens",
      "permission": [],
      "root": "deepseek/deepseek-v4.1-flash",
      "parent": null
    }
  ]
}
FieldMeaning
idThe exact string to send as model. Always provider/model form.
objectAlways "model".
createdUnix timestamp of when the model was added to the catalog.
owned_byAlways "tokens", whichever lab built the model.
root, parent, permissionPresent for OpenAI SDK compatibility; root equals id.

The response does not include prices, context windows or capabilities. Those are on the model catalog, one page per model.

Why a model is missing from the list#

The list is filtered for the key that calls it, so two keys on the same account can see different lists:

  1. Key allow-list. If the key was created with an allowed-models list, only those models appear.
  2. Plan and wallet. Each model is enabled for certain plan tiers and for pay-as-you-go. A model appears if your active plan's tier includes it, or if pay-as-you-go is allowed for it and your wallet balance is above zero.
  3. Catalog status. Only active catalog models are listed.

An empty data array usually means no active plan and no wallet balance. Subscribe or top up in billing; details in plans and wallet.

Check remaining usage with GET /v1/tokens/usage#

This endpoint is specific to Tokens. It returns the current plan, plan usage windows, wallet balance and the calling key's limits. The tokens.mjs usage CLI command reads it, and it's handy in a status bar or a pre-flight check in a long-running agent.

bash
curl https://tokens.bd/v1/tokens/usage \
  -H "Authorization: Bearer $TOKENS_API_KEY"
json
{
  "object": "tokens.usage",
  "plan": {
    "name": "Pro Monthly",
    "tier": "monthly",
    "periodEnd": "2026-11-02T08:15:00.000Z"
  },
  "windows": [
    {
      "type": "session_5h",
      "label": "5-Hour Session",
      "unit": "usd",
      "limit": 5,
      "used": 1.284,
      "remaining": 3.716,
      "percentUsed": 26,
      "resetsAt": "2026-10-03T14:40:00.000Z"
    }
  ],
  "wallet": { "balanceUsd": 12.5 },
  "key": {
    "monthlySpendCapUsd": 20,
    "allowedModels": null
  }
}

The plan name and numbers above are illustrative; yours depend on your plan.

FieldTypeMeaning
planobject or nullActive subscription, or null on pay-as-you-go only
plan.tierstringPlan tier, such as weekly or monthly
plan.periodEndISO 8601 or nullWhen the current subscription period ends
windows[]arrayPlan usage windows that have usage in their current period
windows[].typestringsession_5h (rolling 5 hours), weekly or monthly
windows[].unitstringusd (credits, expressed in USD) or requests
windows[].limit, used, remainingnumberIn the window's unit
windows[].percentUsedinteger0 to 100
windows[].resetsAtISO 8601When the window resets
walletobject or nullbalanceUsd: prepaid wallet balance in USD, or null if there is no wallet
key.monthlySpendCapUsdnumber or nullThe calling key's monthly cap, null if uncapped
key.allowedModelsarray or nullThe key's allow-list, null if any model is allowed

Note

A window that hasn't been used yet in its current period is left out of windows, so an empty array on a fresh plan or right after a reset is normal.

When a window is used up, inference requests return 429 window_exhausted until resetsAt. The dashboard can also notify you at 50, 75, 90 and 100 percent of a window; see usage and alerts.

The response is sent with Cache-Control: no-store. Polling it once a minute is plenty; it counts toward neither your bill nor your per-minute request limit.

Embeddings: POST /v1/embeddings#

The route exists and follows OpenAI's embeddings format, but it only works for catalog models that are embedding models. Chat models will fail on it. Check the model catalog for an embedding model before building on this endpoint; if none is listed, there is no embedding model available for your account yet.

Legacy completions: POST /v1/completions#

The old prompt-in, text-out completions endpoint is passed through for tools that still use it. It only works when the upstream behind a model supports it, and many chat models don't. Use chat completions for anything new.

Endpoints that don't exist#

Any other path under /v1 returns 404 with code unsupported_endpoint. That includes images, audio, files, batches, assistants, fine-tuning and moderations. See errors for the full code list.

Was this page helpful?

Still stuck? Open a support ticket

Need help configuring your agent?

Test your connection with the connection tester, or create an API key.