Skip to content

Connect Warp to Tokens

Point Warp's agents at Tokens with a custom inference endpoint that speaks OpenAI Chat Completions, and know where Warp will not use it.

Works withWarp
On this page

Warp is a terminal with built-in AI agents. It supports a custom inference endpoint: any server that implements OpenAI's POST /v1/chat/completions. Tokens is one, so you can add https://tokens.bd/v1 as an endpoint and pick Tokens models in Warp's model picker.

Limits to know first#

These come from Warp's documentation (checked October 2026):

  • Requests go through Warp's backend. Warp calls your endpoint from its servers, so the endpoint must be a public URL. tokens.bd is public, so this works, but it means your prompts pass through Warp as well as Tokens.
  • "Auto" models and custom routers never use your endpoint. Only a model you pick explicitly goes to Tokens.
  • Cloud Agents don't use it. Warp's docs say the custom endpoint doesn't apply to Cloud Agents.
  • Eligibility. The feature is free for individuals and organizations of 10 or fewer employees. Larger organizations need a Business or Enterprise plan, and those still consume Warp platform credits.

If those limits rule Warp out, a terminal agent that runs entirely on your key, such as Claude Code, OpenCode or Aider, may suit you better.

Add Tokens as a custom inference endpoint in Warp#

1. Create a key#

Create a key in the dashboard. API keys covers spend caps and model allow-lists. A separate key for Warp makes its usage easy to spot.

2. Enter the endpoint#

  1. Open Warp Settings and search for inference endpoint.
  2. Add the endpoint URL: https://tokens.bd/v1. Warp's docs describe this field as "the base URL that exposes /v1/chat/completions", and their example uses a URL ending in /v1.
  3. Add your API key: tok_live_your_key.
  4. Add the model ID: deepseek/deepseek-v4.1-flash.
  5. Save.

In summary:

Warp Settings inference endpoint
Endpoint URL:  https://tokens.bd/v1
API key:       tok_live_your_key
Model ID:      deepseek/deepseek-v4.1-flash

Warp stores the key in its settings and doesn't read TOKENS_API_KEY from your environment.

3. Select the model#

Open Warp's model picker and select the Tokens model explicitly. If the picker is on Auto, requests go to Warp's own models, not Tokens.

Switch models#

Add another model ID in the endpoint settings for each Tokens model you want, then choose between them in the model picker. Use exact IDs, including the provider prefix. Find them in the model catalog or with:

bash
curl -s https://tokens.bd/v1/models -H "Authorization: Bearer $TOKENS_API_KEY"

Warp's agents edit files and run commands, so pick a model with tool calling support. Choosing a model explains the trade-offs.

Verify it works#

Test the key directly:

bash
export TOKENS_API_KEY=tok_live_your_key
curl -s https://tokens.bd/v1/chat/completions \
  -H "Authorization: Bearer $TOKENS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "deepseek/deepseek-v4.1-flash", "max_tokens": 16, "messages": [{"role": "user", "content": "Reply with OK"}]}'

Then, with the Tokens model selected, ask Warp's agent a short question. Open your dashboard usage analytics: the request should appear with the model you chose. If Warp answered but your usage shows nothing, Warp used one of its own models; check the picker.

Troubleshooting#

Nothing reaches Tokens. The picker is on Auto, or you're using a Cloud Agent. Select the Tokens model explicitly, outside Cloud Agents.

The setting is unavailable. Your organization may be above the 10-employee threshold without a Business or Enterprise plan.

401 invalid_api_key. Wrong, revoked or rotated key. Rotation stops the old secret immediately, so update the key in Warp.

404 model_not_found. The model ID must match the Tokens ID exactly, for example deepseek/deepseek-v4.1-flash.

404 on every request. The endpoint URL should be https://tokens.bd/v1. Don't add /chat/completions; Warp appends it.

403 model_not_allowed_on_key. The key's allow-list excludes the model. Allow-lists can't be changed after creation, so create a new key.

429 rate_limited or concurrency_limit. Wait for Retry-After. Details in Troubleshooting.

Warp's requests use the Chat Completions format.

Source: Warp docs, custom inference endpoint, checked October 2026.

Was this page helpful?

Still stuck? Open a support ticket

Need help configuring your agent?

Test your connection with the connection tester, or create an API key.