Open WebUI is a self-hosted chat interface. It talks to any OpenAI-compatible server through an OpenAI connection, which sends Chat Completions requests, so Tokens works with a base URL of https://tokens.bd/v1 and a Tokens key. You add the connection in the admin settings, or with environment variables.
This guide is based on Open WebUI's documentation, checked on 11 October 2026 (the latest release then was v0.12.0). It was checked against the documentation, not run end to end against a live Open WebUI install.
What you need#
- A Tokens key. Create one in the dashboard; API keys covers spend caps and allowed models.
- Admin access to your Open WebUI instance.
- At least one model id from the model catalog. This guide uses
deepseek/deepseek-v4.1-flash. Ids have the formprovider/model.
Add Tokens as a connection#
- In Open WebUI, go to Settings > Admin > Connections.
- In the Manage OpenAI API Connections list, click the add button (Add Connection).
- Leave Connection Type as External.
- Set URL to
https://tokens.bd/v1. - Paste your Tokens key into API Key.
- Click Save.
Enter the base URL only. The examples in Open WebUI's documentation all end at /v1 and none include a path such as /chat/completions, and Tokens' own base URL ends at /v1 too.
Saving does not test the connection. To test it, use Verify Connection, which calls GET /models with a Bearer token. Tokens answers that call with the models your key can use, so a successful check also proves the key works. Leave the Provider setting under Advanced at Default, and leave Forward cookies off. Open WebUI's docs say to enable that only for a server you control that authenticates by cookie, and never for third-party endpoints.
Or set it with environment variables#
Open WebUI reads the connection from environment variables as well. For one connection:
OPENAI_API_BASE_URL=https://tokens.bd/v1
OPENAI_API_KEY=tok_live_your_keyOPENAI_API_BASE_URLS and OPENAI_API_KEYS take several values separated by semicolons. If you use them to add Tokens next to another provider, keep the two lists in the same order. In a Docker Compose file, pass the key from your shell or an .env file next to the compose file instead of typing it into the file you commit:
services:
open-webui:
environment:
- OPENAI_API_BASE_URL=https://tokens.bd/v1
- OPENAI_API_KEY=${TOKENS_API_KEY}These variables are marked as ConfigVar in Open WebUI's reference: the value is stored on first launch, and after that Open WebUI uses the stored value rather than the environment, so later changes are made in Settings > Admin > Connections. The reference describes ENABLE_PERSISTENT_CONFIG=False as the switch that makes environment variables win again, with UI changes lost on restart.
Docker and host.docker.internal
Open WebUI's documentation uses host.docker.internal for model servers that run on the Docker host. Tokens is a remote address, so the URL stays https://tokens.bd/v1.
Choose which models appear#
By default Open WebUI shows every model the connection returns from GET /models. For Tokens that is the list your key is allowed to use, so a key with an allowed-models list shows just those models.
The Model IDs field under Advanced changes this:
- Empty (the default): all models from the provider are detected.
- Filled: the list you type replaces the fetched one. Open WebUI stops calling the provider's
/modelsfor that connection and shows only the ids you entered. Use it to hide models from your users.
Type the exact Tokens id, for example deepseek/deepseek-v4.1-flash, and click the plus button to add it. Duplicates are refused and surrounding spaces are removed.
Ids with a slash#
Tokens ids always contain a slash. Open WebUI's documentation does not describe special handling of slashes in model ids, so the ids are shown as the provider returns them. If you want a label per connection, the Prefix ID field joins your prefix and the model id with a dot (a prefix tokens gives tokens.provider/model) and removes it again before the request goes upstream. Type the prefix without a trailing slash: the docs show that groq/ produces groq/.model.
Check that it works#
- Open a new chat and pick a Tokens model from the model selector.
- Send "Reply with OK".
- The reply streams in, and the request appears in your Tokens usage.
If you want to rule out Open WebUI first, run this from any machine:
curl -s https://tokens.bd/v1/chat/completions \
-H "Authorization: Bearer $TOKENS_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "deepseek/deepseek-v4.1-flash", "max_tokens": 16, "messages": [{"role": "user", "content": "Reply with OK"}]}'Streaming and tool calling#
Open WebUI's docs list POST /v1/chat/completions as the one required endpoint, with streaming and the usual parameters such as temperature, top_p and max_tokens. Tokens streams that endpoint as Server-Sent Events; see Streaming.
Tool use needs a model and a provider that accept tools and tool_choice, which Open WebUI's docs call out as a requirement. Tokens passes these through, but not every model handles them well. Check Choosing a model and the model's page in the catalog before enabling tools for your users, and read Tool calling for the request format.
Background tasks cost credits: set a cheap task model#
Open WebUI sends extra requests behind each conversation. Its task-models page lists them: chat titles, tags, follow-up suggestions, autocomplete, retrieval and web search query rewriting, image prompts and context compaction summaries. Each one is a billed Tokens request.
By default the task model is Current Model, so these requests run on whatever model the user chatted with, an expensive one included. To change that:
- Go to Settings > Admin > Interface and find the Tasks section.
- Set External Task Model to a small, fast, non-reasoning model from your Tokens list. Open WebUI's docs recommend a small, non-reasoning model because reasoning models add latency and cost for simple outputs. This is the picker that applies to Tokens, because it is not a Local connection such as Ollama. The environment variable is
TASK_MODEL_EXTERNAL. - Under Task Model Parameters, set
max_tokensif you want a hard cap on background requests. The docs note that setting any parameter removes the built-in 1000 token limit for titles and summaries, so includemax_tokensyourself. The variable isTASK_MODEL_PARAMS, a JSON object.
If the named task model is unavailable, Open WebUI falls back to the chat's own model instead of failing, which also means a task model that your key may not call silently costs you the expensive model again. Keep the task model in the key's allowed-models list.
To stop tasks altogether, turn off their toggles under Generation in the same admin page: Title Generation (ENABLE_TITLE_GENERATION), Tags Generation (ENABLE_TAGS_GENERATION), Follow Up Generation (ENABLE_FOLLOW_UP_GENERATION) and Autocomplete Generation (ENABLE_AUTOCOMPLETE_GENERATION). Open WebUI's reference lists autocomplete as off by default and the others as on. Autocomplete fires while users type, so leave it off on a metered key. Context compaction has its own picker, Context Compaction Model, under Chat.
Run it on a shared server#
Everyone who uses the instance spends the one key you put in the connection, and shares that account's rate and concurrency limits (Rate limits). Before you give the instance to other people:
- Create a key just for this instance, with a monthly spend cap and an allowed-models list (API keys).
- Put the task model in that list, as described above.
- Keep the key out of files you commit. Use the environment variable or the admin page.
- Watch usage in the dashboard; Open WebUI does not show you what the Tokens account has left.
GET /v1/tokens/usagedoes (Models and usage).
What does not go through this connection#
- Images, speech and transcription. Open WebUI has separate settings and variables for image generation, text-to-speech and speech-to-text. Tokens does not serve those endpoints; they return 404
unsupported_endpoint. Point those features at another provider. - Embeddings for RAG. Open WebUI's docs mark
/v1/embeddingsas optional, used for RAG. Tokens serves embeddings only for models that are embedding models (Embeddings). Keep the default RAG embedding setup unless you pick one of those.
Troubleshooting#
Verify Connection fails, or the model list is empty. Check the URL ends in /v1 and has no trailing path. An empty list from Tokens usually means no active plan and no wallet balance; see Models and usage. Open WebUI's docs also say a failed check does not always mean chat is broken: add the model id under Model IDs and try a chat.
The settings page hangs or the list loads slowly. Open WebUI times out the model list fetch after 10 seconds by default; raise AIOHTTP_CLIENT_TIMEOUT_MODEL_LIST on slow networks. Open WebUI's connection-error guide has a section on model list loading problems for a saved URL that cannot be reached.
401 invalid_api_key or missing_api_key. The key is wrong, rotated or empty. Paste a fresh one into the connection. Rotation stops the old secret immediately.
404 model_not_found. The model id is not the full Tokens id. Copy it from the model list, provider/model included.
403 model_not_allowed_on_key. The key's allowed-models list does not include the model. This also happens when a task model is outside the list; chat titles fail while normal chat works.
402 insufficient_credits, 403 monthly_spend_cap_exceeded. Out of credits, or the key hit its cap. Top up in billing or use a key with a higher cap.
429 rate_limited or concurrency_limit. Many users, or many background tasks, on one key. Wait for Retry-After, turn off the task toggles you don't need, or use a second key for a second instance.
Every code is listed in Errors; more fixes are in Troubleshooting. If you ask support for help, include the x-tokens-request-id header from a curl run with -i.
Sources: Open WebUI documentation, OpenAI-compatible providers, Starting with OpenAI, task models and the environment variable reference, checked 11 October 2026.