# LlamaIndex

> Use Tokens from LlamaIndex (Python) with the OpenAILike class: api_base, is_chat_model, context window, streaming, function-calling agents, and embeddings for RAG.

LlamaIndex is a Python framework for building RAG pipelines and agents over your own data. It talks to models through LLM classes, and `OpenAILike` is the one meant for any OpenAI-compatible server. Pointed at `https://tokens.bd/v1`, it sends Chat Completions requests to Tokens. A RAG pipeline also needs an embedding model, and LlamaIndex picks an OpenAI one unless you say otherwise, so that part gets its own section below.

:::note[Checked against the documentation]
Based on the LlamaIndex documentation and the integration packages' source and README files, checked October 2026 against `llama-index-llms-openai-like` 0.8.1 (released 1 October 2026, needs `llama-index-core` 0.14.3 or newer and Python 3.10 or newer). The code was checked against the documentation, not run end to end against Tokens.
:::

## What you need

- A Tokens key from [API keys](/docs/api-keys), exported as `TOKENS_API_KEY`.
- A model id from [/models](/models), for example `deepseek/deepseek-v4.1-flash`, and its context window from the model's page.
- For RAG, an embedding model id from the catalog (see [Embeddings](#embeddings-for-rag)).

```bash
pip install --upgrade llama-index-llms-openai-like
export TOKENS_API_KEY="tok_live_your_key"
```

## Set up OpenAILike

```python title="chat.py"
import os

from llama_index.core.llms import ChatMessage
from llama_index.llms.openai_like import OpenAILike

llm = OpenAILike(
    model="deepseek/deepseek-v4.1-flash",
    api_base="https://tokens.bd/v1",
    api_key=os.environ["TOKENS_API_KEY"],
    is_chat_model=True,
    context_window=128000,  # set this to the window shown on the model's page in /models
)

response = llm.chat(
    [
        ChatMessage(role="system", content="You are a precise SQL reviewer."),
        ChatMessage(role="user", content="Is SELECT * in a view a problem? Two sentences."),
    ]
)
print(response)
```

The arguments, from the integration's README:

| Argument                    | What to set                                                                                                                             |
| --------------------------- | --------------------------------------------------------------------------------------------------------------------------------------- |
| `model`                     | The Tokens model id, exactly as listed in [/models](/models) or by `GET https://tokens.bd/v1/models`. LlamaIndex sends it unchanged.     |
| `api_base`                  | `https://tokens.bd/v1`. Keep the `/v1`.                                                                                                  |
| `api_key`                   | Your Tokens key.                                                                                                                        |
| `is_chat_model`             | `True`. See below.                                                                                                                      |
| `context_window`            | The model's real context window in tokens. LlamaIndex uses it to size prompts and chunks.                                               |
| `is_function_calling_model` | `True` if the model supports tool calling and you use agents or tools. The default is `False`.                                          |

Two defaults catch people out:

- **`is_chat_model` defaults to `False`.** In that mode `OpenAILike` sends your calls to the completions endpoint, `/v1/completions`. Many chat models do not serve it and the call fails ([Legacy completions](/docs/legacy-completions)). Set `is_chat_model=True` so requests go to `/v1/chat/completions`.
- **`context_window` has a small default** (about 3,900 tokens, according to the class docstring). Leave it and LlamaIndex splits and truncates context far more than needed. Take the real value from the model's page.

`OpenAILike` inherits its timeout (60 seconds) and retries (3) from LlamaIndex's `OpenAI` class. Both can be set in the constructor as `timeout` and `max_retries`. For long generations raise `timeout`, or stream. Tokens already fails over between upstream sources before it answers ([Rate limits](/docs/rate-limits)), so a few retries are enough.

### Make it the default for a pipeline

```python
from llama_index.core import Settings

Settings.llm = llm
```

Everything that uses the default LLM (query engines, chat engines) now goes to Tokens. You can also pass the model to one engine: `index.as_query_engine(llm=llm)`.

## Stream responses

```python
messages = [ChatMessage(role="user", content="Give me three naming tips for Python modules.")]

for chunk in llm.stream_chat(messages):
    print(chunk.delta, end="", flush=True)
print()
```

`delta` is the new text in each chunk. Tokens' streamed responses carry token counts only when the request asks for them ([Streaming](/docs/streaming)), and the source does not show `OpenAILike` asking for them, so streamed calls may show no usage in LlamaIndex. The dashboard still records the request.

## Tool calling and agents

Set `is_function_calling_model=True` and give LlamaIndex's `FunctionAgent` plain Python functions. Their names, type hints and docstrings become the tool schema.

```python title="agent.py"
import asyncio
import os

from llama_index.core.agent.workflow import AgentStream, FunctionAgent
from llama_index.llms.openai_like import OpenAILike

llm = OpenAILike(
    model="deepseek/deepseek-v4.1-flash",
    api_base="https://tokens.bd/v1",
    api_key=os.environ["TOKENS_API_KEY"],
    is_chat_model=True,
    is_function_calling_model=True,
    context_window=128000,  # set this to the window shown on the model's page in /models
)


def get_weather(city: str) -> str:
    """Current weather for a city."""
    return f"{city}: light rain, 29C"  # replace with a real lookup


agent = FunctionAgent(
    tools=[get_weather],
    llm=llm,
    system_prompt="You answer questions about the weather. Use the tool when you need data.",
)


async def main() -> None:
    response = await agent.run(user_msg="Do I need an umbrella in Dhaka?")
    print(response)

    # The same agent, with streamed text:
    handler = agent.run(user_msg="And in Chattogram?")
    async for event in handler.stream_events():
        if isinstance(event, AgentStream):
            print(event.delta, end="", flush=True)
    await handler
    print()


asyncio.run(main())
```

`agent.run(...)` returns a handler. `await` it for the final answer, or iterate `handler.stream_events()` first and then `await handler`, as in the LlamaIndex streaming guide. `AgentStream.delta` is the newest piece of text. If your model cannot stream a reply that includes tool calls, create the agent with `streaming=False`, which the LlamaIndex documentation gives for that case.

The agent calls the model once per step, so one question that needs two tool calls is at least three requests to Tokens, and they count against your [rate limits](/docs/rate-limits). Tool calling needs a model that supports it ([Tool calling](/docs/tool-calling)).
`should_use_structured_outputs=True` on `OpenAILike` switches structured output on through `response_format`. Only set it for a model that supports JSON schema output ([Structured output](/docs/structured-output)).

## Embeddings for RAG

If you set `Settings.llm` and nothing else, `VectorStoreIndex` still embeds with LlamaIndex's default, OpenAI's `text-embedding-ada-002` through `OpenAIEmbedding`. That call needs an OpenAI key and goes to OpenAI, not Tokens. You have two choices.

**Embed through Tokens.** `/v1/embeddings` works only with embedding models. A chat model such as `deepseek/deepseek-v4.1-flash` is not one, and the call fails. Find an embedding model in the catalog (search for `embed` in [/models](/models)) and read its id from an environment variable, as the [Embeddings](/docs/embeddings) page does, because the catalog changes. If the catalog has no embedding model, none is available to your account yet, so use the second choice.

```bash
pip install --upgrade llama-index-embeddings-openai-like
export EMBEDDING_MODEL="the-id-from-the-catalog"
```

```python
import os

from llama_index.core import Settings
from llama_index.embeddings.openai_like import OpenAILikeEmbedding

Settings.embed_model = OpenAILikeEmbedding(
    model_name=os.environ["EMBEDDING_MODEL"],
    api_base="https://tokens.bd/v1",
    api_key=os.environ["TOKENS_API_KEY"],
    embed_batch_size=10,
)
```

The class takes `model_name`, not `model`, and passing `model` raises a `ValueError`. `embed_batch_size` is how many texts go in one request. Larger batches mean fewer requests, within the provider's limit for that model. Not documented by LlamaIndex: which `encoding_format` it requests. If calls fail or return odd vectors, read the base64 note on the [Embeddings](/docs/embeddings) page, since the OpenAI client library defaults to base64 and not every provider supports it. Do not mix vectors from different models in one index, and re-embed if you change models.

**Embed somewhere else.** Set another embedding model, for example a local one, in `Settings.embed_model` (the LlamaIndex documentation lists the options) and use Tokens only for the LLM. The query step then still sends your retrieved text to Tokens as part of the prompt, like any other chat request.

## Check that it works

Run `python chat.py`. A short answer printed to the terminal means the key, base URL and model id are right. Then open Usage analytics in the [dashboard](/dashboard): the request should be listed under the model id you set. To list the ids your key can use, run `curl -H "Authorization: Bearer $TOKENS_API_KEY" https://tokens.bd/v1/models` ([cURL](/docs/curl)).

## Choosing a model

For plain question answering over retrieved text, most chat models work. Agents and structured extraction need reliable tool calls or JSON output, which differs by model. [Choosing a model](/docs/choosing-a-model) covers which suit which job, and each model's page in [/models](/models) lists its context window. A bigger `context_window` lets LlamaIndex put more retrieved chunks in a prompt, which also costs more input tokens per question.

## Limits and what does not work

- **Model names.** `OpenAILike` sends whatever `model` you give it, so the id must be one Tokens serves. Nothing in LlamaIndex checks it against the catalog.
- **Embeddings from chat models.** Not possible, see above.
- **OpenAI-only integrations.** Features of LlamaIndex's own `OpenAI` class, such as the Responses API classes, assume OpenAI. Use `OpenAILike` for Tokens.
- **Token counting.** Tokens bills on the usage the provider reports, not on LlamaIndex's own counts ([Token counting](/docs/token-counting)).
- **Hosted tools.** Tools run by OpenAI itself (hosted web search, file search) do not exist on Tokens. Use your own Python tools.
- **Other LlamaIndex LLM classes.** This page covers only `OpenAILike`. LlamaIndex's Anthropic class is not documented for custom gateways, so it is not covered.

## Troubleshooting

**A 400 or 404 on `/v1/completions`.** `is_chat_model` is still `False`, so the call goes to the completions endpoint, which many chat models do not serve. Set `is_chat_model=True`.

**404 on every request.** `api_base` is missing the `/v1`, or has it twice. Use `https://tokens.bd/v1` exactly.

**404 `model_not_found`.** The id has a typo, or it is an embedding model used for chat (or the reverse). Check it against `GET https://tokens.bd/v1/models`.

**An OpenAI error about a missing or wrong key during indexing.** Embeddings are going to OpenAI, since `Settings.embed_model` is still the default. Set it as shown above.

**The index fails with 400 `invalid_request` or 404 `model_not_found` on `/v1/embeddings`.** The id you gave `OpenAILikeEmbedding` is not an embedding model. See [Embeddings](/docs/embeddings).

**The agent answers without calling the tool, or errors about tool calling.** Check `is_function_calling_model=True` and that the model supports tools ([Tool calling](/docs/tool-calling)). Try another model.

**Context errors or very short answers.** `context_window` is the default. Set it to the model's real window.

**401 `invalid_api_key`, 402 `insufficient_credits`, 403 `model_not_allowed_on_key`, 429 `window_exhausted` or `concurrency_limit`.** These are account limits, not LlamaIndex problems. Indexing many documents in parallel can reach `concurrency_limit`. Lower the parallelism or the number of workers. See [Errors](/docs/errors) and [Troubleshooting](/docs/troubleshooting), and include the `x-tokens-request-id` response header when you contact [support](/docs/support).

---
Page: https://tokens.bd/docs/llamaindex
