# Open WebUI

> Add Tokens as an OpenAI connection in Open WebUI, list your models, move background tasks such as chat titles to a cheap model, and fix connection errors.

Open WebUI is a self-hosted chat interface. It talks to any OpenAI-compatible server through an OpenAI connection, which sends Chat Completions requests, so Tokens works with a base URL of `https://tokens.bd/v1` and a Tokens key. You add the connection in the admin settings, or with environment variables.

This guide is based on Open WebUI's documentation, checked on 11 October 2026 (the latest release then was v0.12.0). It was checked against the documentation, not run end to end against a live Open WebUI install.

## What you need

- A Tokens key. Create one in the dashboard; [API keys](/docs/api-keys) covers spend caps and allowed models.
- Admin access to your Open WebUI instance.
- At least one model id from the [model catalog](/models). This guide uses `deepseek/deepseek-v4.1-flash`. Ids have the form `provider/model`.

## Add Tokens as a connection

1. In Open WebUI, go to **Settings > Admin > Connections**.
2. In the **Manage OpenAI API Connections** list, click the add button (**Add Connection**).
3. Leave **Connection Type** as **External**.
4. Set **URL** to `https://tokens.bd/v1`.
5. Paste your Tokens key into **API Key**.
6. Click **Save**.

Enter the base URL only. The examples in Open WebUI's documentation all end at `/v1` and none include a path such as `/chat/completions`, and Tokens' own base URL ends at `/v1` too.

Saving does not test the connection. To test it, use **Verify Connection**, which calls `GET /models` with a Bearer token. Tokens answers that call with the models your key can use, so a successful check also proves the key works. Leave the **Provider** setting under **Advanced** at **Default**, and leave **Forward cookies** off. Open WebUI's docs say to enable that only for a server you control that authenticates by cookie, and never for third-party endpoints.

### Or set it with environment variables

Open WebUI reads the connection from environment variables as well. For one connection:

```bash title="Environment of the Open WebUI container"
OPENAI_API_BASE_URL=https://tokens.bd/v1
OPENAI_API_KEY=tok_live_your_key
```

`OPENAI_API_BASE_URLS` and `OPENAI_API_KEYS` take several values separated by semicolons. If you use them to add Tokens next to another provider, keep the two lists in the same order. In a Docker Compose file, pass the key from your shell or an `.env` file next to the compose file instead of typing it into the file you commit:

```yaml title="docker-compose.yml (excerpt)"
services:
  open-webui:
    environment:
      - OPENAI_API_BASE_URL=https://tokens.bd/v1
      - OPENAI_API_KEY=${TOKENS_API_KEY}
```

These variables are marked as `ConfigVar` in Open WebUI's reference: the value is stored on first launch, and after that Open WebUI uses the stored value rather than the environment, so later changes are made in **Settings > Admin > Connections**. The reference describes `ENABLE_PERSISTENT_CONFIG=False` as the switch that makes environment variables win again, with UI changes lost on restart.

:::note[Docker and host.docker.internal]
Open WebUI's documentation uses `host.docker.internal` for model servers that run on the Docker host. Tokens is a remote address, so the URL stays `https://tokens.bd/v1`.
:::

## Choose which models appear

By default Open WebUI shows every model the connection returns from `GET /models`. For Tokens that is the list your key is allowed to use, so a key with an allowed-models list shows just those models.

The **Model IDs** field under **Advanced** changes this:

- Empty (the default): all models from the provider are detected.
- Filled: the list you type replaces the fetched one. Open WebUI stops calling the provider's `/models` for that connection and shows only the ids you entered. Use it to hide models from your users.

Type the exact Tokens id, for example `deepseek/deepseek-v4.1-flash`, and click the plus button to add it. Duplicates are refused and surrounding spaces are removed.

### Ids with a slash

Tokens ids always contain a slash. Open WebUI's documentation does not describe special handling of slashes in model ids, so the ids are shown as the provider returns them. If you want a label per connection, the **Prefix ID** field joins your prefix and the model id with a dot (a prefix `tokens` gives `tokens.provider/model`) and removes it again before the request goes upstream. Type the prefix without a trailing slash: the docs show that `groq/` produces `groq/.model`.

## Check that it works

1. Open a new chat and pick a Tokens model from the model selector.
2. Send "Reply with OK".
3. The reply streams in, and the request appears in your Tokens [usage](/dashboard/usage).

If you want to rule out Open WebUI first, run this from any machine:

```bash
curl -s https://tokens.bd/v1/chat/completions \
  -H "Authorization: Bearer $TOKENS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "deepseek/deepseek-v4.1-flash", "max_tokens": 16, "messages": [{"role": "user", "content": "Reply with OK"}]}'
```

## Streaming and tool calling

Open WebUI's docs list `POST /v1/chat/completions` as the one required endpoint, with streaming and the usual parameters such as temperature, top_p and max_tokens. Tokens streams that endpoint as Server-Sent Events; see [Streaming](/docs/streaming).

Tool use needs a model and a provider that accept `tools` and `tool_choice`, which Open WebUI's docs call out as a requirement. Tokens passes these through, but not every model handles them well. Check [Choosing a model](/docs/choosing-a-model) and the model's page in the [catalog](/models) before enabling tools for your users, and read [Tool calling](/docs/tool-calling) for the request format.

## Background tasks cost credits: set a cheap task model

Open WebUI sends extra requests behind each conversation. Its task-models page lists them: chat titles, tags, follow-up suggestions, autocomplete, retrieval and web search query rewriting, image prompts and context compaction summaries. Each one is a billed Tokens request.

By default the task model is **Current Model**, so these requests run on whatever model the user chatted with, an expensive one included. To change that:

1. Go to **Settings > Admin > Interface** and find the **Tasks** section.
2. Set **External Task Model** to a small, fast, non-reasoning model from your Tokens list. Open WebUI's docs recommend a small, non-reasoning model because reasoning models add latency and cost for simple outputs. This is the picker that applies to Tokens, because it is not a Local connection such as Ollama. The environment variable is `TASK_MODEL_EXTERNAL`.
3. Under **Task Model Parameters**, set `max_tokens` if you want a hard cap on background requests. The docs note that setting any parameter removes the built-in 1000 token limit for titles and summaries, so include `max_tokens` yourself. The variable is `TASK_MODEL_PARAMS`, a JSON object.

If the named task model is unavailable, Open WebUI falls back to the chat's own model instead of failing, which also means a task model that your key may not call silently costs you the expensive model again. Keep the task model in the key's allowed-models list.

To stop tasks altogether, turn off their toggles under **Generation** in the same admin page: Title Generation (`ENABLE_TITLE_GENERATION`), Tags Generation (`ENABLE_TAGS_GENERATION`), Follow Up Generation (`ENABLE_FOLLOW_UP_GENERATION`) and Autocomplete Generation (`ENABLE_AUTOCOMPLETE_GENERATION`). Open WebUI's reference lists autocomplete as off by default and the others as on. Autocomplete fires while users type, so leave it off on a metered key. Context compaction has its own picker, **Context Compaction Model**, under **Chat**.

## Run it on a shared server

Everyone who uses the instance spends the one key you put in the connection, and shares that account's rate and concurrency limits ([Rate limits](/docs/rate-limits)). Before you give the instance to other people:

- Create a key just for this instance, with a monthly spend cap and an allowed-models list ([API keys](/docs/api-keys)).
- Put the task model in that list, as described above.
- Keep the key out of files you commit. Use the environment variable or the admin page.
- Watch usage in the dashboard; Open WebUI does not show you what the Tokens account has left. `GET /v1/tokens/usage` does ([Models and usage](/docs/models-and-usage)).

## What does not go through this connection

- **Images, speech and transcription.** Open WebUI has separate settings and variables for image generation, text-to-speech and speech-to-text. Tokens does not serve those endpoints; they return 404 `unsupported_endpoint`. Point those features at another provider.
- **Embeddings for RAG.** Open WebUI's docs mark `/v1/embeddings` as optional, used for RAG. Tokens serves embeddings only for models that are embedding models ([Embeddings](/docs/embeddings)). Keep the default RAG embedding setup unless you pick one of those.

## Troubleshooting

**Verify Connection fails, or the model list is empty.** Check the URL ends in `/v1` and has no trailing path. An empty list from Tokens usually means no active plan and no wallet balance; see [Models and usage](/docs/models-and-usage). Open WebUI's docs also say a failed check does not always mean chat is broken: add the model id under **Model IDs** and try a chat.

**The settings page hangs or the list loads slowly.** Open WebUI times out the model list fetch after 10 seconds by default; raise `AIOHTTP_CLIENT_TIMEOUT_MODEL_LIST` on slow networks. Open WebUI's connection-error guide has a section on model list loading problems for a saved URL that cannot be reached.

**401 `invalid_api_key` or `missing_api_key`.** The key is wrong, rotated or empty. Paste a fresh one into the connection. Rotation stops the old secret immediately.

**404 `model_not_found`.** The model id is not the full Tokens id. Copy it from the model list, `provider/model` included.

**403 `model_not_allowed_on_key`.** The key's allowed-models list does not include the model. This also happens when a task model is outside the list; chat titles fail while normal chat works.

**402 `insufficient_credits`, 403 `monthly_spend_cap_exceeded`.** Out of credits, or the key hit its cap. Top up in [billing](/dashboard/billing) or use a key with a higher cap.

**429 `rate_limited` or `concurrency_limit`.** Many users, or many background tasks, on one key. Wait for `Retry-After`, turn off the task toggles you don't need, or use a second key for a second instance.

Every code is listed in [Errors](/docs/errors); more fixes are in [Troubleshooting](/docs/troubleshooting). If you ask [support](/docs/support) for help, include the `x-tokens-request-id` header from a curl run with `-i`.

Sources: Open WebUI documentation, [OpenAI-compatible providers](https://docs.openwebui.com/getting-started/quick-start/connect-a-provider/starting-with-openai-compatible/), [Starting with OpenAI](https://docs.openwebui.com/getting-started/quick-start/connect-a-provider/starting-with-openai), [task models](https://docs.openwebui.com/features/administration/task-models/) and the [environment variable reference](https://docs.openwebui.com/reference/env-configuration), checked 11 October 2026.

---
Page: https://tokens.bd/docs/open-webui
