Skip to content

LlamaIndex

Use Tokens from LlamaIndex (Python) with the OpenAILike class: api_base, is_chat_model, context window, streaming, function-calling agents, and embeddings for RAG.

Works withLlamaIndex
On this page

LlamaIndex is a Python framework for building RAG pipelines and agents over your own data. It talks to models through LLM classes, and OpenAILike is the one meant for any OpenAI-compatible server. Pointed at https://tokens.bd/v1, it sends Chat Completions requests to Tokens. A RAG pipeline also needs an embedding model, and LlamaIndex picks an OpenAI one unless you say otherwise, so that part gets its own section below.

Checked against the documentation

Based on the LlamaIndex documentation and the integration packages' source and README files, checked October 2026 against llama-index-llms-openai-like 0.8.1 (released 1 October 2026, needs llama-index-core 0.14.3 or newer and Python 3.10 or newer). The code was checked against the documentation, not run end to end against Tokens.

What you need#

  • A Tokens key from API keys, exported as TOKENS_API_KEY.
  • A model id from /models, for example deepseek/deepseek-v4.1-flash, and its context window from the model's page.
  • For RAG, an embedding model id from the catalog (see Embeddings).
bash
pip install --upgrade llama-index-llms-openai-like
export TOKENS_API_KEY="tok_live_your_key"

Set up OpenAILike#

chat.py
import os

from llama_index.core.llms import ChatMessage
from llama_index.llms.openai_like import OpenAILike

llm = OpenAILike(
    model="deepseek/deepseek-v4.1-flash",
    api_base="https://tokens.bd/v1",
    api_key=os.environ["TOKENS_API_KEY"],
    is_chat_model=True,
    context_window=128000,  # set this to the window shown on the model's page in /models
)

response = llm.chat(
    [
        ChatMessage(role="system", content="You are a precise SQL reviewer."),
        ChatMessage(role="user", content="Is SELECT * in a view a problem? Two sentences."),
    ]
)
print(response)

The arguments, from the integration's README:

ArgumentWhat to set
modelThe Tokens model id, exactly as listed in /models or by GET https://tokens.bd/v1/models. LlamaIndex sends it unchanged.
api_basehttps://tokens.bd/v1. Keep the /v1.
api_keyYour Tokens key.
is_chat_modelTrue. See below.
context_windowThe model's real context window in tokens. LlamaIndex uses it to size prompts and chunks.
is_function_calling_modelTrue if the model supports tool calling and you use agents or tools. The default is False.

Two defaults catch people out:

  • is_chat_model defaults to False. In that mode OpenAILike sends your calls to the completions endpoint, /v1/completions. Many chat models do not serve it and the call fails (Legacy completions). Set is_chat_model=True so requests go to /v1/chat/completions.
  • context_window has a small default (about 3,900 tokens, according to the class docstring). Leave it and LlamaIndex splits and truncates context far more than needed. Take the real value from the model's page.

OpenAILike inherits its timeout (60 seconds) and retries (3) from LlamaIndex's OpenAI class. Both can be set in the constructor as timeout and max_retries. For long generations raise timeout, or stream. Tokens already fails over between upstream sources before it answers (Rate limits), so a few retries are enough.

Make it the default for a pipeline#

python
from llama_index.core import Settings

Settings.llm = llm

Everything that uses the default LLM (query engines, chat engines) now goes to Tokens. You can also pass the model to one engine: index.as_query_engine(llm=llm).

Stream responses#

python
messages = [ChatMessage(role="user", content="Give me three naming tips for Python modules.")]

for chunk in llm.stream_chat(messages):
    print(chunk.delta, end="", flush=True)
print()

delta is the new text in each chunk. Tokens' streamed responses carry token counts only when the request asks for them (Streaming), and the source does not show OpenAILike asking for them, so streamed calls may show no usage in LlamaIndex. The dashboard still records the request.

Tool calling and agents#

Set is_function_calling_model=True and give LlamaIndex's FunctionAgent plain Python functions. Their names, type hints and docstrings become the tool schema.

agent.py
import asyncio
import os

from llama_index.core.agent.workflow import AgentStream, FunctionAgent
from llama_index.llms.openai_like import OpenAILike

llm = OpenAILike(
    model="deepseek/deepseek-v4.1-flash",
    api_base="https://tokens.bd/v1",
    api_key=os.environ["TOKENS_API_KEY"],
    is_chat_model=True,
    is_function_calling_model=True,
    context_window=128000,  # set this to the window shown on the model's page in /models
)


def get_weather(city: str) -> str:
    """Current weather for a city."""
    return f"{city}: light rain, 29C"  # replace with a real lookup


agent = FunctionAgent(
    tools=[get_weather],
    llm=llm,
    system_prompt="You answer questions about the weather. Use the tool when you need data.",
)


async def main() -> None:
    response = await agent.run(user_msg="Do I need an umbrella in Dhaka?")
    print(response)

    # The same agent, with streamed text:
    handler = agent.run(user_msg="And in Chattogram?")
    async for event in handler.stream_events():
        if isinstance(event, AgentStream):
            print(event.delta, end="", flush=True)
    await handler
    print()


asyncio.run(main())

agent.run(...) returns a handler. await it for the final answer, or iterate handler.stream_events() first and then await handler, as in the LlamaIndex streaming guide. AgentStream.delta is the newest piece of text. If your model cannot stream a reply that includes tool calls, create the agent with streaming=False, which the LlamaIndex documentation gives for that case.

The agent calls the model once per step, so one question that needs two tool calls is at least three requests to Tokens, and they count against your rate limits. Tool calling needs a model that supports it (Tool calling). should_use_structured_outputs=True on OpenAILike switches structured output on through response_format. Only set it for a model that supports JSON schema output (Structured output).

Embeddings for RAG#

If you set Settings.llm and nothing else, VectorStoreIndex still embeds with LlamaIndex's default, OpenAI's text-embedding-ada-002 through OpenAIEmbedding. That call needs an OpenAI key and goes to OpenAI, not Tokens. You have two choices.

Embed through Tokens. /v1/embeddings works only with embedding models. A chat model such as deepseek/deepseek-v4.1-flash is not one, and the call fails. Find an embedding model in the catalog (search for embed in /models) and read its id from an environment variable, as the Embeddings page does, because the catalog changes. If the catalog has no embedding model, none is available to your account yet, so use the second choice.

bash
pip install --upgrade llama-index-embeddings-openai-like
export EMBEDDING_MODEL="the-id-from-the-catalog"
python
import os

from llama_index.core import Settings
from llama_index.embeddings.openai_like import OpenAILikeEmbedding

Settings.embed_model = OpenAILikeEmbedding(
    model_name=os.environ["EMBEDDING_MODEL"],
    api_base="https://tokens.bd/v1",
    api_key=os.environ["TOKENS_API_KEY"],
    embed_batch_size=10,
)

The class takes model_name, not model, and passing model raises a ValueError. embed_batch_size is how many texts go in one request. Larger batches mean fewer requests, within the provider's limit for that model. Not documented by LlamaIndex: which encoding_format it requests. If calls fail or return odd vectors, read the base64 note on the Embeddings page, since the OpenAI client library defaults to base64 and not every provider supports it. Do not mix vectors from different models in one index, and re-embed if you change models.

Embed somewhere else. Set another embedding model, for example a local one, in Settings.embed_model (the LlamaIndex documentation lists the options) and use Tokens only for the LLM. The query step then still sends your retrieved text to Tokens as part of the prompt, like any other chat request.

Check that it works#

Run python chat.py. A short answer printed to the terminal means the key, base URL and model id are right. Then open Usage analytics in the dashboard: the request should be listed under the model id you set. To list the ids your key can use, run curl -H "Authorization: Bearer $TOKENS_API_KEY" https://tokens.bd/v1/models (cURL).

Choosing a model#

For plain question answering over retrieved text, most chat models work. Agents and structured extraction need reliable tool calls or JSON output, which differs by model. Choosing a model covers which suit which job, and each model's page in /models lists its context window. A bigger context_window lets LlamaIndex put more retrieved chunks in a prompt, which also costs more input tokens per question.

Limits and what does not work#

  • Model names. OpenAILike sends whatever model you give it, so the id must be one Tokens serves. Nothing in LlamaIndex checks it against the catalog.
  • Embeddings from chat models. Not possible, see above.
  • OpenAI-only integrations. Features of LlamaIndex's own OpenAI class, such as the Responses API classes, assume OpenAI. Use OpenAILike for Tokens.
  • Token counting. Tokens bills on the usage the provider reports, not on LlamaIndex's own counts (Token counting).
  • Hosted tools. Tools run by OpenAI itself (hosted web search, file search) do not exist on Tokens. Use your own Python tools.
  • Other LlamaIndex LLM classes. This page covers only OpenAILike. LlamaIndex's Anthropic class is not documented for custom gateways, so it is not covered.

Troubleshooting#

A 400 or 404 on /v1/completions. is_chat_model is still False, so the call goes to the completions endpoint, which many chat models do not serve. Set is_chat_model=True.

404 on every request. api_base is missing the /v1, or has it twice. Use https://tokens.bd/v1 exactly.

404 model_not_found. The id has a typo, or it is an embedding model used for chat (or the reverse). Check it against GET https://tokens.bd/v1/models.

An OpenAI error about a missing or wrong key during indexing. Embeddings are going to OpenAI, since Settings.embed_model is still the default. Set it as shown above.

The index fails with 400 invalid_request or 404 model_not_found on /v1/embeddings. The id you gave OpenAILikeEmbedding is not an embedding model. See Embeddings.

The agent answers without calling the tool, or errors about tool calling. Check is_function_calling_model=True and that the model supports tools (Tool calling). Try another model.

Context errors or very short answers. context_window is the default. Set it to the model's real window.

401 invalid_api_key, 402 insufficient_credits, 403 model_not_allowed_on_key, 429 window_exhausted or concurrency_limit. These are account limits, not LlamaIndex problems. Indexing many documents in parallel can reach concurrency_limit. Lower the parallelism or the number of workers. See Errors and Troubleshooting, and include the x-tokens-request-id response header when you contact support.

Was this page helpful?

Still stuck? Open a support ticket

Need help configuring your agent?

Test your connection with the connection tester, or create an API key.