Skip to content

LangChain and LiteLLM

Use Tokens from LangChain (Python and JavaScript) with ChatOpenAI and a custom base URL, and from LiteLLM as an SDK or proxy with the openai/ model prefix.

Works withLangChain PythonLangChain JSLiteLLM
On this page

LangChain and LiteLLM both have an OpenAI client that takes a custom base URL, and that's all Tokens needs. This page shows the LangChain ChatOpenAI setup in Python and JavaScript, then LiteLLM as a Python SDK and as a proxy.

All examples assume your key is exported:

bash
export TOKENS_API_KEY="tok_live_your_key"

LangChain Python: ChatOpenAI with base_url#

bash
pip install --upgrade langchain-openai
chain.py
import os
from langchain_openai import ChatOpenAI

llm = ChatOpenAI(
    model="deepseek/deepseek-v4.1-flash",
    base_url="https://tokens.bd/v1",
    api_key=os.environ["TOKENS_API_KEY"],
    temperature=0.2,
    max_tokens=500,
    timeout=120,
    max_retries=2,
)

reply = llm.invoke([
    ("system", "You are a precise SQL reviewer."),
    ("human", "Is SELECT * in a view a problem? Two sentences."),
])
print(reply.content)
print(reply.usage_metadata)

model takes the Tokens model ID exactly as listed in /models or GET /v1/models. LangChain passes it through unchanged.

Streaming#

python
llm = ChatOpenAI(
    model="deepseek/deepseek-v4.1-flash",
    base_url="https://tokens.bd/v1",
    api_key=os.environ["TOKENS_API_KEY"],
    stream_usage=True,  # sets stream_options.include_usage so you get token counts
)

for chunk in llm.stream("Give me three naming tips for Python modules."):
    print(chunk.content, end="", flush=True)

Without stream_usage=True, streamed responses from Tokens carry no usage data. See Streaming.

Tools and chains#

bind_tools, with_structured_output and LCEL chains all work, since they only depend on the OpenAI chat format. Tool calling and structured output still need a model that supports them. Check the model page in /models before building on one, and see Tool calling.

python
from langchain_core.tools import tool

@tool
def get_weather(city: str) -> str:
    """Current weather for a city."""
    return f"{city}: light rain, 29C"

llm_with_tools = llm.bind_tools([get_weather])
msg = llm_with_tools.invoke("Do I need an umbrella in Dhaka?")
print(msg.tool_calls)

LangChain JS: ChatOpenAI with configuration.baseURL#

bash
npm install @langchain/openai @langchain/core

In JavaScript the base URL goes inside configuration, which is passed straight to the underlying OpenAI client.

chain.ts
import { ChatOpenAI } from "@langchain/openai";

const llm = new ChatOpenAI({
  model: "deepseek/deepseek-v4.1-flash",
  apiKey: process.env.TOKENS_API_KEY,
  configuration: { baseURL: "https://tokens.bd/v1" },
  temperature: 0.2,
  maxTokens: 500,
  streamUsage: true,
});

const reply = await llm.invoke("Explain debouncing vs throttling in two sentences.");
console.log(reply.content);

for await (const chunk of await llm.stream("List three uses for a WeakMap.")) {
  process.stdout.write(String(chunk.content));
}

Run LangChain code on a server or in a CLI, not in browser bundles. Tokens doesn't send CORS headers, and the key would be visible to anyone who opens dev tools.

LiteLLM: use the openai/ prefix with api_base#

LiteLLM routes a request by the prefix of the model name. openai/ tells it to use its generic OpenAI-compatible client against whatever api_base you give it. LiteLLM strips that first openai/, so the Tokens model ID (which has its own provider/ part) is sent as-is.

LiteLLM Python SDK#

bash
pip install --upgrade litellm
litellm_example.py
import os
import litellm

response = litellm.completion(
    model="openai/deepseek/deepseek-v4.1-flash",  # "openai/" + Tokens model ID
    api_base="https://tokens.bd/v1",
    api_key=os.environ["TOKENS_API_KEY"],
    messages=[{"role": "user", "content": "One tip for faster pytest runs?"}],
    max_tokens=200,
)
print(response.choices[0].message.content)

Keep /v1 on api_base. Leaving it off is the most common cause of 404s with LiteLLM.

LiteLLM proxy config#

If your team runs a LiteLLM proxy, add Tokens models to model_list and keep the key in the environment:

config.yaml
model_list:
  - model_name: deepseek-v4.1-flash # the name your apps will request
    litellm_params:
      model: openai/deepseek/deepseek-v4.1-flash
      api_base: https://tokens.bd/v1
      api_key: os.environ/TOKENS_API_KEY
bash
litellm --config config.yaml

Clients of the proxy then ask for deepseek-v4.1-flash, and LiteLLM forwards those requests to Tokens. A few things to know:

  • Spend, plan windows and rate limits are enforced per Tokens account, not per proxy user. Everyone behind one Tokens key shares that account's per-minute and concurrency limits (Rate limits).
  • If you want separate ceilings for different apps, create separate Tokens keys, each with a monthly spend cap and an allowed-models list (API keys), and use one key per model_list entry.
  • LiteLLM's own cost tracking won't know Tokens prices. Use the usage page in the dashboard or GET /v1/tokens/usage as the source of truth.

When something fails#

LangChain and LiteLLM both wrap the OpenAI SDK's errors, so the HTTP status and the Tokens error.code (invalid_api_key, model_not_found, insufficient_credits, window_exhausted, and so on) end up in the exception message. The Troubleshooting guide maps each code to a fix. When you open a ticket, include the x-tokens-request-id response header. To get at it, reproduce the call once with cURL and -i.

Was this page helpful?

Still stuck? Open a support ticket

Need help configuring your agent?

Test your connection with the connection tester, or create an API key.