LangChain and LiteLLM both have an OpenAI client that takes a custom base URL, and that's all Tokens needs. This page shows the LangChain ChatOpenAI setup in Python and JavaScript, then LiteLLM as a Python SDK and as a proxy.
All examples assume your key is exported:
export TOKENS_API_KEY="tok_live_your_key"LangChain Python: ChatOpenAI with base_url#
pip install --upgrade langchain-openaiimport os
from langchain_openai import ChatOpenAI
llm = ChatOpenAI(
model="deepseek/deepseek-v4.1-flash",
base_url="https://tokens.bd/v1",
api_key=os.environ["TOKENS_API_KEY"],
temperature=0.2,
max_tokens=500,
timeout=120,
max_retries=2,
)
reply = llm.invoke([
("system", "You are a precise SQL reviewer."),
("human", "Is SELECT * in a view a problem? Two sentences."),
])
print(reply.content)
print(reply.usage_metadata)model takes the Tokens model ID exactly as listed in /models or GET /v1/models. LangChain passes it through unchanged.
Streaming#
llm = ChatOpenAI(
model="deepseek/deepseek-v4.1-flash",
base_url="https://tokens.bd/v1",
api_key=os.environ["TOKENS_API_KEY"],
stream_usage=True, # sets stream_options.include_usage so you get token counts
)
for chunk in llm.stream("Give me three naming tips for Python modules."):
print(chunk.content, end="", flush=True)Without stream_usage=True, streamed responses from Tokens carry no usage data. See Streaming.
Tools and chains#
bind_tools, with_structured_output and LCEL chains all work, since they only depend on the OpenAI chat format. Tool calling and structured output still need a model that supports them. Check the model page in /models before building on one, and see Tool calling.
from langchain_core.tools import tool
@tool
def get_weather(city: str) -> str:
"""Current weather for a city."""
return f"{city}: light rain, 29C"
llm_with_tools = llm.bind_tools([get_weather])
msg = llm_with_tools.invoke("Do I need an umbrella in Dhaka?")
print(msg.tool_calls)LangChain JS: ChatOpenAI with configuration.baseURL#
npm install @langchain/openai @langchain/coreIn JavaScript the base URL goes inside configuration, which is passed straight to the underlying OpenAI client.
import { ChatOpenAI } from "@langchain/openai";
const llm = new ChatOpenAI({
model: "deepseek/deepseek-v4.1-flash",
apiKey: process.env.TOKENS_API_KEY,
configuration: { baseURL: "https://tokens.bd/v1" },
temperature: 0.2,
maxTokens: 500,
streamUsage: true,
});
const reply = await llm.invoke("Explain debouncing vs throttling in two sentences.");
console.log(reply.content);
for await (const chunk of await llm.stream("List three uses for a WeakMap.")) {
process.stdout.write(String(chunk.content));
}Run LangChain code on a server or in a CLI, not in browser bundles. Tokens doesn't send CORS headers, and the key would be visible to anyone who opens dev tools.
LiteLLM: use the openai/ prefix with api_base#
LiteLLM routes a request by the prefix of the model name. openai/ tells it to use its generic OpenAI-compatible client against whatever api_base you give it. LiteLLM strips that first openai/, so the Tokens model ID (which has its own provider/ part) is sent as-is.
LiteLLM Python SDK#
pip install --upgrade litellmimport os
import litellm
response = litellm.completion(
model="openai/deepseek/deepseek-v4.1-flash", # "openai/" + Tokens model ID
api_base="https://tokens.bd/v1",
api_key=os.environ["TOKENS_API_KEY"],
messages=[{"role": "user", "content": "One tip for faster pytest runs?"}],
max_tokens=200,
)
print(response.choices[0].message.content)Keep /v1 on api_base. Leaving it off is the most common cause of 404s with LiteLLM.
LiteLLM proxy config#
If your team runs a LiteLLM proxy, add Tokens models to model_list and keep the key in the environment:
model_list:
- model_name: deepseek-v4.1-flash # the name your apps will request
litellm_params:
model: openai/deepseek/deepseek-v4.1-flash
api_base: https://tokens.bd/v1
api_key: os.environ/TOKENS_API_KEYlitellm --config config.yamlClients of the proxy then ask for deepseek-v4.1-flash, and LiteLLM forwards those requests to Tokens. A few things to know:
- Spend, plan windows and rate limits are enforced per Tokens account, not per proxy user. Everyone behind one Tokens key shares that account's per-minute and concurrency limits (Rate limits).
- If you want separate ceilings for different apps, create separate Tokens keys, each with a monthly spend cap and an allowed-models list (API keys), and use one key per
model_listentry. - LiteLLM's own cost tracking won't know Tokens prices. Use the usage page in the dashboard or
GET /v1/tokens/usageas the source of truth.
When something fails#
LangChain and LiteLLM both wrap the OpenAI SDK's errors, so the HTTP status and the Tokens error.code (invalid_api_key, model_not_found, insufficient_credits, window_exhausted, and so on) end up in the exception message. The Troubleshooting guide maps each code to a fix. When you open a ticket, include the x-tokens-request-id response header. To get at it, reproduce the call once with cURL and -i.