# CrewAI

> Run CrewAI agents and crews on Tokens: the LLM class with a custom base URL and custom_openai, how to write the model id, streaming, tool calls, rate limits and troubleshooting.

CrewAI is a Python framework for agents that work together as a crew. Each agent gets an `LLM`, and CrewAI can send that LLM's requests to any OpenAI-compatible endpoint. Tokens is one: CrewAI sends OpenAI Chat Completions requests to `https://tokens.bd/v1` with your Tokens key.

The one detail that trips people up is the model string. Tokens model ids already contain a slash (`provider/model`), and CrewAI also reads the part before the first slash. The next section shows what to type.

:::note[What was checked]
Based on CrewAI's LLMs documentation (docs.crewai.com, checked October 2026) and the `crewai` 1.15.27 source on PyPI and GitHub (released 9 October 2026). The routing rules below come from `llm.py` in that source, because the docs page does not describe them. The code was checked against the documentation and source, not run end to end against Tokens.
:::

## What you need

- A Tokens key from [API keys](/docs/api-keys), exported as `TOKENS_API_KEY`.
- A model id from [/models](/models). Pick one that supports tool calling if your agents use tools ([Choosing a model](/docs/choosing-a-model)).
- Python 3.10 to 3.13 (CrewAI 1.15.27 declares `>=3.10,<3.14`).

```bash
pip install crewai
export TOKENS_API_KEY="tok_live_your_key"
```

CrewAI's docs install it with `uv` (`uv tool install crewai` for the command line tool, `uv add crewai` inside a project). `pip` works the same way for a plain script. `crewai` already depends on the `openai` package, so you do not need the `litellm` extra for this setup.

## Set up the LLM

```python title="llm.py"
import os
from crewai import LLM

llm = LLM(
    model="deepseek/deepseek-v4.1-flash",
    base_url="https://tokens.bd/v1",
    api_key=os.environ["TOKENS_API_KEY"],
    custom_openai=True,
    timeout=120,
    max_retries=2,
)
```

What each part does:

| Setting | Why |
| --- | --- |
| `model` | The Tokens model id, exactly as listed in [/models](/models). |
| `base_url` | `https://tokens.bd/v1`, including `/v1`. CrewAI hands it to the OpenAI Python SDK, which appends `/chat/completions`. |
| `custom_openai=True` | Forces CrewAI's native OpenAI client for this LLM, whatever the model string looks like. CrewAI's own gateway example uses it. It also needs a custom endpoint, so it raises an error if you forget `base_url`. |
| `timeout`, `max_retries` | Seconds to wait for a response and retry count. CrewAI's docs list both. If you leave them out, the OpenAI SDK defaults apply (10 minutes, 2 retries). |

### How to write the model id

With `custom_openai=True`, type the Tokens id exactly as it appears in the catalog, with no extra prefix. CrewAI removes a leading `openai/` if there is one and sends the rest unchanged, so `deepseek/deepseek-v4.1-flash` reaches Tokens as `deepseek/deepseek-v4.1-flash`.

Two cases to watch:

- **Do not drop `custom_openai=True`.** Without it, CrewAI reads the part before the first slash as a provider name. A Tokens id can start with a name that CrewAI treats as one of its own providers (`deepseek/` is one), and the request then goes to that provider's client instead of Tokens. The LiteLLM-style form, `model="openai/<tokens id>"` plus `base_url`, also reaches the OpenAI client, but `custom_openai=True` removes the guesswork.
- **Ids that start with `openai/`.** The leading `openai/` is stripped. If a Tokens id itself begins with `openai/`, write it twice: `model="openai/openai/<name>"`.

CrewAI's LLMs page says to always include a provider prefix. The Tokens id already is `provider/model`, so that rule is met, and its own custom-endpoint example passes the gateway's id this way.

### Use environment variables instead

CrewAI also reads `OPENAI_BASE_URL` and `OPENAI_API_KEY` (its docs list both):

```bash
export OPENAI_BASE_URL="https://tokens.bd/v1"
export OPENAI_API_KEY="$TOKENS_API_KEY"
```

```python
from crewai import LLM

llm = LLM(model="deepseek/deepseek-v4.1-flash", custom_openai=True)
```

Those variables are global. Any other OpenAI-based tool in the same shell will also send its traffic to Tokens. Passing `base_url` and `api_key` in code keeps the change local to CrewAI.

## Check that it works

First call the LLM on its own, with no agents:

```python title="check.py"
from llm import llm

print(llm.call("Reply with the single word: ready"))
```

A normal sentence back means the key, the base URL and the model id are all right. To check the id without spending tokens, list the catalog:

```bash
curl -s https://tokens.bd/v1/models -H "Authorization: Bearer $TOKENS_API_KEY"
```

## A minimal crew

```python title="crew.py"
from crewai import Agent, Crew, Task
from llm import llm

reviewer = Agent(
    role="Python code reviewer",
    goal="Point out bugs and unclear names in short snippets",
    backstory="You review pull requests for a small team and keep feedback short.",
    llm=llm,
    max_iter=5,
)

task = Task(
    description="Review this function and list the problems:\n\ndef avg(xs): return sum(xs)/len(xs)",
    expected_output="A bullet list with at most three items.",
    agent=reviewer,
)

crew = Crew(agents=[reviewer], tasks=[task])
result = crew.kickoff()

print(result.raw)
print(crew.usage_metrics)
```

`result.raw` is the final text. `crew.usage_metrics` reports token counts for the run. Your Tokens bill comes from the dashboard's [usage page](/docs/usage-and-alerts), not from CrewAI's numbers.

`max_iter` caps how many reasoning steps an agent may take before it must answer (CrewAI's default is 20). Each step is one request that Tokens bills, so a lower cap is cheaper while you are still testing.

## Stream responses

Set `stream=True` on the LLM:

```python
streaming_llm = LLM(
    model="deepseek/deepseek-v4.1-flash",
    base_url="https://tokens.bd/v1",
    api_key=os.environ["TOKENS_API_KEY"],
    custom_openai=True,
    stream=True,
)
```

CrewAI emits an `LLMStreamChunkEvent` for each chunk. This listener is the example from CrewAI's LLMs page:

```python
from crewai.events import BaseEventListener, LLMStreamChunkEvent

class MyCustomListener(BaseEventListener):
    def setup_listeners(self, crewai_event_bus):
        @crewai_event_bus.on(LLMStreamChunkEvent)
        def on_llm_stream_chunk(self, event: LLMStreamChunkEvent):
            print(f"Received chunk: {event.chunk}")

my_listener = MyCustomListener()
```

Create the listener before you run the crew. When streaming is on, CrewAI asks Tokens for `stream_options.include_usage`, so the usage numbers still arrive. See [Streaming](/docs/streaming) for the wire format.

## Tool calls

Give an agent tools with `@tool` from `crewai.tools` and the `tools` argument:

```python
from crewai import Agent
from crewai.tools import tool
from llm import llm

@tool("Get weather")
def get_weather(city: str) -> str:
    """Current weather for a city. Use it when the user asks about the weather."""
    return f"{city}: light rain, 29C"

agent = Agent(
    role="Travel helper",
    goal="Answer weather questions for travellers",
    backstory="You know Bangladesh well.",
    llm=llm,
    tools=[get_weather],
)
```

There is no function-calling switch to set for a custom endpoint. In CrewAI's source, the OpenAI provider's `supports_function_calling()` returns true unless the model is an o1-style model, so CrewAI sends tools in the OpenAI format. Whether the model then calls them well depends on the model: check its page in [/models](/models) and read [Tool calling](/docs/tool-calling). If a model ignores tools, try another one before changing code.

## Limits and what to watch

- **Requests per minute.** Crews with several agents can fire many requests in a short time. Set `max_rpm` on the `Agent` (CrewAI's docs describe it as the cap on requests per minute) to stay under your plan's limit. Going over returns `429 rate_limited`. See [Rate limits](/docs/rate-limits).
- **Parallel tasks.** Tasks that run at the same time count against your account's concurrent request limit and can return `429 concurrency_limit`. Run them in sequence or lower the parallelism.
- **Long generations.** The gateway waits up to 600 seconds for a response. Keep CrewAI's `timeout` at or below that, and use `stream=True` for very long outputs.
- **Not covered here.** CrewAI features that call an embeddings model on their own (memory, knowledge sources) are not set up by this page. Tokens serves embeddings ([Embeddings](/docs/embeddings)), but you have to point those features at the same endpoint yourself.

## Troubleshooting

| Symptom | Cause and fix |
| --- | --- |
| `ImportError: Unable to initialize LLM ... LiteLLM fallback package is not installed` | CrewAI did not choose its OpenAI client and fell back to LiteLLM. Add `custom_openai=True` and `base_url`. Installing `crewai[litellm]` also silences it, but then LiteLLM does the routing. |
| The error mentions another provider's API key | The model string matched one of CrewAI's own providers. Add `custom_openai=True`. |
| `401 missing_api_key` or `invalid_api_key` | The key did not reach the request. Pass `api_key=` explicitly and check `TOKENS_API_KEY` is set in the process that runs the crew. See [API keys](/docs/api-keys). |
| `404 model_not_found` | The id is wrong, or CrewAI stripped an `openai/` that belonged to the id. Copy the id from [/models](/models); for ids that start with `openai/`, write the prefix twice. |
| `402 insufficient_credits` | The plan credits and wallet cannot cover the request. Top up in [billing](/dashboard/billing). Loops that run up to `max_iter` steps use credits fast. |
| `429 rate_limited` or `concurrency_limit` | Too many requests. Set `max_rpm`, run tasks in sequence, wait the `Retry-After` seconds. `window_exhausted` means a plan window is used up, so retrying will not help until it resets. |
| Timeouts | Raise `timeout` on the `LLM`, or stream. A `504 upstream_timeout` comes from the upstream after 600 seconds. Try a smaller request. |

Every code is in [Errors](/docs/errors). When you ask [support](/docs/support) for help, include the `x-tokens-request-id` response header. The examples on this page do not expose response headers, so reproduce the call once with [cURL](/docs/curl) and `-i` to read it.

---
Page: https://tokens.bd/docs/crewai
