Cline is an open-source autonomous coding agent. It runs as a VS Code extension and also ships a CLI and a desktop app. Cline connects to Tokens through its OpenAI Compatible provider, which speaks the OpenAI Chat Completions protocol to https://tokens.bd/v1.
Before you start#
You need:
- Cline installed. In VS Code, open the Extensions view and search for "Cline".
- A Tokens key. Create one in the dashboard; API keys explains spend caps and model allow-lists. Coding agents send many requests per task, so a monthly spend cap on this key is cheap insurance.
- A model ID. This guide uses
deepseek/deepseek-v4.1-flash. Browse others in the model catalog.
Configure the OpenAI Compatible provider in Cline#
Open the Cline panel, then its settings, and fill in these fields:
| Field | Value |
|---|---|
| API Provider | OpenAI Compatible |
| Base URL | https://tokens.bd/v1 |
| API Key | tok_live_your_key |
| Model ID | deepseek/deepseek-v4.1-flash |
| Use Azure Identity Authentication | Off |
The same values, with your real key filled in, are on Connect Your Agent.
Cline keeps the key in its own settings field. It does not read a TOKENS_API_KEY environment variable, so paste the key directly.
Model Configuration#
Under Model Configuration, Cline asks for details it can't discover from the endpoint:
- Context Window size. Set the model's real context window from the model catalog. Cline uses this to decide when the conversation is too long. A number that's too high leads to context-length errors from the upstream; too low wastes capacity.
- Max Output Tokens. The model's output limit, or a lower number if you want shorter replies.
- Computer Use. This is Cline's tool use switch. Cline is an agent that edits files and runs commands through tool calls, so leave it on for models that support tool calling.
- Image Support. Turn on only if the model accepts image input. The catalog lists this per model.
- Input/Output Price. Cline uses these only to show its own cost estimate in the panel. Tokens bills from its own usage meter, and your dashboard is the number that counts. You can copy prices from the catalog to make Cline's estimate closer, or leave them at zero.
Switch models#
Change the Model ID field to another Tokens ID and update the Model Configuration values to match the new model. The context window in particular differs a lot between models.
To see which IDs your key can use:
curl -s https://tokens.bd/v1/models -H "Authorization: Bearer $TOKENS_API_KEY"If you're choosing between fast, cheap models and stronger reasoning models for agent work, Choosing a model has the reasoning.
Verify it works#
First confirm the key and model outside the editor:
export TOKENS_API_KEY=tok_live_your_key
curl -s https://tokens.bd/v1/chat/completions \
-H "Authorization: Bearer $TOKENS_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "deepseek/deepseek-v4.1-flash", "max_tokens": 16, "messages": [{"role": "user", "content": "Reply with OK"}]}'Then send a simple chat message in the Cline panel, such as "List the files in this folder." A working setup answers and, for that prompt, makes a tool call. The request appears in your dashboard usage analytics shortly after.
Troubleshooting#
Cline's own troubleshooting points at three things: the key, the model ID, and the base URL. Here is what each looks like with Tokens.
401 invalid_api_key. The key is wrong, revoked, or was rotated. Rotation stops the old secret immediately, so paste the new one into Cline.
"Model Not Found" or 404 model_not_found. The Model ID must match exactly, including the provider prefix. deepseek-v4.1-flash without deepseek/ will not resolve.
404 or connection errors on every request. Check the Base URL. It must be https://tokens.bd/v1 with /v1 and no trailing /chat/completions.
403 model_not_allowed_on_key. The key was created with an allow-list that doesn't include this model. Allow-lists can't be edited after creation, so create a new key.
429 rate_limited or concurrency_limit. Agents can burst past the per-minute limit during long tasks. The response carries Retry-After; Cline retries, or you can wait and resume. A 429 window_exhausted means your plan's usage window is used up until it resets.
402 insufficient_credits. Add funds or renew in billing.
The agent stalls or loops without editing files. The model may handle tool calls poorly. Try a model that the catalog lists with tool support.
More error codes are in Troubleshooting, and the request shape is in Chat Completions. For a terminal agent instead of an editor extension, see OpenCode or Claude Code.
Source: Cline docs, OpenAI Compatible provider, checked October 2026.