LibreChat is a self-hosted chat application. It adds OpenAI-compatible providers as custom endpoints in a file called librechat.yaml. A custom endpoint with the base URL https://tokens.bd/v1 sends Chat Completions requests to Tokens. LibreChat can also talk to the Anthropic Messages API through a custom endpoint, covered near the end.
This guide is based on LibreChat's documentation, checked on 11 October 2026. The latest release then was v0.8.8, and the librechat.example.yaml in its repository uses config version: 1.3.17. It was checked against the documentation, not run end to end against a live LibreChat install.
What you need#
- A Tokens key. Create one in the dashboard; API keys covers spend caps and allowed models.
- A LibreChat install where you can edit
librechat.yamland.env, and restart it. - A model id from the model catalog. This guide uses
deepseek/deepseek-v4.1-flash. Ids have the formprovider/model.
Add the key to .env#
LibreChat's documentation recommends referencing keys from .env with ${VAR_NAME} instead of writing them into librechat.yaml. Add the key to the .env file in the LibreChat project root:
TOKENS_API_KEY=tok_live_your_keyKeep .env out of git. Every ${...} reference in the config needs a matching entry. If one is missing, the endpoint still appears, and the error Missing API Key for <endpoint> shows up when a message is sent.
Add the endpoint to librechat.yaml#
Create librechat.yaml next to .env (or set CONFIG_PATH in .env to its full path) and add Tokens under endpoints.custom:
version: 1.3.17
cache: true
endpoints:
custom:
- name: 'Tokens'
apiKey: '${TOKENS_API_KEY}'
baseURL: 'https://tokens.bd/v1'
models:
default: ['deepseek/deepseek-v4.1-flash']
fetch: true
titleConvo: true
titleModel: 'deepseek/deepseek-v4.1-flash'
modelDisplayLabel: 'Tokens'If you already have a librechat.yaml, add only the - name: 'Tokens' block under the existing endpoints.custom list and keep your own version.
What each field does, from LibreChat's reference:
name: required, unique (compared case-insensitively). It is the title in the endpoint selector.apiKeyandbaseURL: required.${TOKENS_API_KEY}reads the value from.env. LibreChat also acceptsuser_provided, where each user types their own key in the interface. Writing the key into the file as plain text is possible but the docs advise against it.models: required.defaultis a non-empty list and is the fallback if fetching fails.fetch: truemakes LibreChat ask the endpoint for its model list.titleConvoandtitleModel: turn on conversation titles and choose the model that writes them. See below, because the default is not a Tokens model.modelDisplayLabel: the name shown next to responses.
LibreChat silently drops an entry if name, baseURL, apiKey or models is missing or misspelled, or if models has neither fetch: true nor a default list. A schema error anywhere in the file stops the server and disables all custom endpoints, so check docker compose logs api if LibreChat does not come up.
Docker: mount the file#
A Docker container cannot see librechat.yaml unless you mount it. Copy docker-compose.override.yml.example to docker-compose.override.yml, then uncomment the bind mount under the api service that maps ./librechat.yaml to /app/librechat.yaml. Compose merges the override file with docker-compose.yml automatically.
After any change to librechat.yaml or .env, restart:
docker compose down && docker compose up -dFor a local install without Docker, stop the process and run npm run backend. Then open LibreChat and check that Tokens is in the endpoint selector.
Models and the slash in ids#
With fetch: true, LibreChat calls the endpoint's model list when it needs it. Tokens answers GET https://tokens.bd/v1/models with only the models your key can use, so a key with an allowed-models list shows just those. LibreChat's docs warn that fetching can slow first use if the endpoint responds late. default stays as the fallback.
To pin a fixed list instead, set fetch: false and list the ids you want:
models:
default: ['deepseek/deepseek-v4.1-flash', 'provider/another-model']
fetch: falseReplace provider/another-model with an id from the catalog.
LibreChat's documentation does not describe special handling of slashes in model ids. Its own OpenRouter example, whose ids look like meta-llama/llama-3-70b-instruct, uses them as plain strings in default and titleModel, and the same shape applies here: write the full Tokens id, provider/model.
Check that it works#
- Select Tokens in the endpoint selector and pick a model.
- Send "Reply with OK". The response streams in, and the request shows in your Tokens usage.
To test the key outside LibreChat:
curl -s https://tokens.bd/v1/chat/completions \
-H "Authorization: Bearer $TOKENS_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "deepseek/deepseek-v4.1-flash", "max_tokens": 16, "messages": [{"role": "user", "content": "Reply with OK"}]}'Streaming and tool calling#
Chat responses stream as Server-Sent Events, which Tokens passes through (Streaming). LibreChat's custom-endpoint page does not cover agents or tool calling, so this guide does not describe how LibreChat uses tools with a custom endpoint. If you use a feature that sends tool definitions, the model needs to support tool calling; see Tool calling and Choosing a model.
If Tokens returns 400 invalid_request because the upstream rejected a request parameter, LibreChat has dropParams to remove a default parameter from requests (for example dropParams: ['stop']) and addParams to add or override one. Use them only after you have seen which parameter is rejected.
Chat titles cost credits: set a cheap titleModel#
When titleConvo is true, LibreChat sends an extra request to name each conversation. Two settings matter, both from LibreChat's reference:
titleConvodefaults tofalse. Leave it that way to send no title requests.titleModeldefaults togpt-3.5-turbo. That is not a Tokens model, so withtitleConvo: trueand notitleModel, the title request fails with 404model_not_found.
Set titleModel to a small, fast Tokens model, not the model the user is chatting with:
titleConvo: true
titleModel: 'provider/cheap-model-id'Replace it with a cheap model from the catalog. The value current_model is also accepted and uses the conversation's own model, which is the expensive choice when users chat with a large model.
Keep that model in the key's allowed-models list. Otherwise titles fail with 403 model_not_allowed_on_key while normal chat works.
Use the Anthropic Messages API instead#
For models you want to reach through /v1/messages, LibreChat supports a custom endpoint with provider: 'anthropic'. Its docs say to use that only for endpoints that implement the native Anthropic Messages API, and to omit provider for OpenAI-compatible gateways such as the entry above. Tokens serves both. The baseURL for this provider is the API root, not the /v1/messages path, so it is https://tokens.bd:
- name: 'Tokens Messages'
provider: 'anthropic'
apiKey: '${TOKENS_API_KEY}'
baseURL: 'https://tokens.bd'
headers:
anthropic-version: '2023-06-01'
models:
default: ['deepseek/deepseek-v4.1-flash']
fetch: false
titleConvo: true
titleModel: 'provider/cheap-model-id'For this provider models.fetch is not used, so every model must be listed under models.default. The anthropic-version header comes from LibreChat's own Anthropic example. Tokens accepts the key in x-api-key or as a Bearer token (Authentication); the format is in Messages. Most people only need the first entry.
Run it on a shared server#
All users of one LibreChat instance spend one Tokens account if you put a single key in .env, and share its rate and concurrency limits (Rate limits). Create a key just for the instance with a monthly spend cap and an allowed-models list (API keys), including your titleModel in the list. The other option, apiKey: 'user_provided', lets each user enter their own Tokens key in the interface, so each person's usage and cap stay separate. LibreChat stores those keys encrypted.
What does not go through Tokens#
Tokens serves chat. Image, audio and file endpoints return 404 unsupported_endpoint, and embeddings work only for embedding models (Models and usage). Keep LibreChat's features that need those on another provider.
Troubleshooting#
Tokens is missing from the endpoint selector. LibreChat drops entries without a clear error. Check name, baseURL, apiKey and models for typos, that the block is indented under endpoints.custom, that no other endpoint has the same name, and that Docker mounts librechat.yaml. Restart after every change.
LibreChat does not start. A schema error in the file disables custom endpoints and stops the server. Read docker compose logs api.
Missing API Key for Tokens, or 401 missing_api_key / invalid_api_key. TOKENS_API_KEY is not in .env or the container was not restarted, or the key is wrong or rotated. Rotation stops the old secret immediately.
404 model_not_found. The id is not a full Tokens id, or the title model is still the gpt-3.5-turbo default. Check ids with GET https://tokens.bd/v1/models.
403 model_not_allowed_on_key. The model, often the titleModel, is not in the key's allowed list.
402 insufficient_credits, 403 monthly_spend_cap_exceeded. Out of credits or the key hit its cap. See billing.
429 rate_limited or concurrency_limit. Too many users or requests on one key; wait for Retry-After. Titles add one request per new chat.
404 on every request. baseURL must be https://tokens.bd/v1 with /v1 and nothing after it, unless you set directEndpoint: true, which LibreChat's reference describes as using baseURL as the full completions URL.
Every code is listed in Errors, with more fixes in Troubleshooting.
Sources: LibreChat documentation, Custom Endpoints quick start, librechat.yaml and the custom endpoint object reference, plus the LibreChat repository's librechat.example.yaml, checked 11 October 2026.