Skip to content

CrewAI

Run CrewAI agents and crews on Tokens: the LLM class with a custom base URL and custom_openai, how to write the model id, streaming, tool calls, rate limits and troubleshooting.

Works withCrewAI
On this page

CrewAI is a Python framework for agents that work together as a crew. Each agent gets an LLM, and CrewAI can send that LLM's requests to any OpenAI-compatible endpoint. Tokens is one: CrewAI sends OpenAI Chat Completions requests to https://tokens.bd/v1 with your Tokens key.

The one detail that trips people up is the model string. Tokens model ids already contain a slash (provider/model), and CrewAI also reads the part before the first slash. The next section shows what to type.

What was checked

Based on CrewAI's LLMs documentation (docs.crewai.com, checked October 2026) and the crewai 1.15.27 source on PyPI and GitHub (released 9 October 2026). The routing rules below come from llm.py in that source, because the docs page does not describe them. The code was checked against the documentation and source, not run end to end against Tokens.

What you need#

  • A Tokens key from API keys, exported as TOKENS_API_KEY.
  • A model id from /models. Pick one that supports tool calling if your agents use tools (Choosing a model).
  • Python 3.10 to 3.13 (CrewAI 1.15.27 declares >=3.10,<3.14).
bash
pip install crewai
export TOKENS_API_KEY="tok_live_your_key"

CrewAI's docs install it with uv (uv tool install crewai for the command line tool, uv add crewai inside a project). pip works the same way for a plain script. crewai already depends on the openai package, so you do not need the litellm extra for this setup.

Set up the LLM#

llm.py
import os
from crewai import LLM

llm = LLM(
    model="deepseek/deepseek-v4.1-flash",
    base_url="https://tokens.bd/v1",
    api_key=os.environ["TOKENS_API_KEY"],
    custom_openai=True,
    timeout=120,
    max_retries=2,
)

What each part does:

SettingWhy
modelThe Tokens model id, exactly as listed in /models.
base_urlhttps://tokens.bd/v1, including /v1. CrewAI hands it to the OpenAI Python SDK, which appends /chat/completions.
custom_openai=TrueForces CrewAI's native OpenAI client for this LLM, whatever the model string looks like. CrewAI's own gateway example uses it. It also needs a custom endpoint, so it raises an error if you forget base_url.
timeout, max_retriesSeconds to wait for a response and retry count. CrewAI's docs list both. If you leave them out, the OpenAI SDK defaults apply (10 minutes, 2 retries).

How to write the model id#

With custom_openai=True, type the Tokens id exactly as it appears in the catalog, with no extra prefix. CrewAI removes a leading openai/ if there is one and sends the rest unchanged, so deepseek/deepseek-v4.1-flash reaches Tokens as deepseek/deepseek-v4.1-flash.

Two cases to watch:

  • Do not drop custom_openai=True. Without it, CrewAI reads the part before the first slash as a provider name. A Tokens id can start with a name that CrewAI treats as one of its own providers (deepseek/ is one), and the request then goes to that provider's client instead of Tokens. The LiteLLM-style form, model="openai/<tokens id>" plus base_url, also reaches the OpenAI client, but custom_openai=True removes the guesswork.
  • Ids that start with openai/. The leading openai/ is stripped. If a Tokens id itself begins with openai/, write it twice: model="openai/openai/<name>".

CrewAI's LLMs page says to always include a provider prefix. The Tokens id already is provider/model, so that rule is met, and its own custom-endpoint example passes the gateway's id this way.

Use environment variables instead#

CrewAI also reads OPENAI_BASE_URL and OPENAI_API_KEY (its docs list both):

bash
export OPENAI_BASE_URL="https://tokens.bd/v1"
export OPENAI_API_KEY="$TOKENS_API_KEY"
python
from crewai import LLM

llm = LLM(model="deepseek/deepseek-v4.1-flash", custom_openai=True)

Those variables are global. Any other OpenAI-based tool in the same shell will also send its traffic to Tokens. Passing base_url and api_key in code keeps the change local to CrewAI.

Check that it works#

First call the LLM on its own, with no agents:

check.py
from llm import llm

print(llm.call("Reply with the single word: ready"))

A normal sentence back means the key, the base URL and the model id are all right. To check the id without spending tokens, list the catalog:

bash
curl -s https://tokens.bd/v1/models -H "Authorization: Bearer $TOKENS_API_KEY"

A minimal crew#

crew.py
from crewai import Agent, Crew, Task
from llm import llm

reviewer = Agent(
    role="Python code reviewer",
    goal="Point out bugs and unclear names in short snippets",
    backstory="You review pull requests for a small team and keep feedback short.",
    llm=llm,
    max_iter=5,
)

task = Task(
    description="Review this function and list the problems:\n\ndef avg(xs): return sum(xs)/len(xs)",
    expected_output="A bullet list with at most three items.",
    agent=reviewer,
)

crew = Crew(agents=[reviewer], tasks=[task])
result = crew.kickoff()

print(result.raw)
print(crew.usage_metrics)

result.raw is the final text. crew.usage_metrics reports token counts for the run. Your Tokens bill comes from the dashboard's usage page, not from CrewAI's numbers.

max_iter caps how many reasoning steps an agent may take before it must answer (CrewAI's default is 20). Each step is one request that Tokens bills, so a lower cap is cheaper while you are still testing.

Stream responses#

Set stream=True on the LLM:

python
streaming_llm = LLM(
    model="deepseek/deepseek-v4.1-flash",
    base_url="https://tokens.bd/v1",
    api_key=os.environ["TOKENS_API_KEY"],
    custom_openai=True,
    stream=True,
)

CrewAI emits an LLMStreamChunkEvent for each chunk. This listener is the example from CrewAI's LLMs page:

python
from crewai.events import BaseEventListener, LLMStreamChunkEvent

class MyCustomListener(BaseEventListener):
    def setup_listeners(self, crewai_event_bus):
        @crewai_event_bus.on(LLMStreamChunkEvent)
        def on_llm_stream_chunk(self, event: LLMStreamChunkEvent):
            print(f"Received chunk: {event.chunk}")

my_listener = MyCustomListener()

Create the listener before you run the crew. When streaming is on, CrewAI asks Tokens for stream_options.include_usage, so the usage numbers still arrive. See Streaming for the wire format.

Tool calls#

Give an agent tools with @tool from crewai.tools and the tools argument:

python
from crewai import Agent
from crewai.tools import tool
from llm import llm

@tool("Get weather")
def get_weather(city: str) -> str:
    """Current weather for a city. Use it when the user asks about the weather."""
    return f"{city}: light rain, 29C"

agent = Agent(
    role="Travel helper",
    goal="Answer weather questions for travellers",
    backstory="You know Bangladesh well.",
    llm=llm,
    tools=[get_weather],
)

There is no function-calling switch to set for a custom endpoint. In CrewAI's source, the OpenAI provider's supports_function_calling() returns true unless the model is an o1-style model, so CrewAI sends tools in the OpenAI format. Whether the model then calls them well depends on the model: check its page in /models and read Tool calling. If a model ignores tools, try another one before changing code.

Limits and what to watch#

  • Requests per minute. Crews with several agents can fire many requests in a short time. Set max_rpm on the Agent (CrewAI's docs describe it as the cap on requests per minute) to stay under your plan's limit. Going over returns 429 rate_limited. See Rate limits.
  • Parallel tasks. Tasks that run at the same time count against your account's concurrent request limit and can return 429 concurrency_limit. Run them in sequence or lower the parallelism.
  • Long generations. The gateway waits up to 600 seconds for a response. Keep CrewAI's timeout at or below that, and use stream=True for very long outputs.
  • Not covered here. CrewAI features that call an embeddings model on their own (memory, knowledge sources) are not set up by this page. Tokens serves embeddings (Embeddings), but you have to point those features at the same endpoint yourself.

Troubleshooting#

SymptomCause and fix
ImportError: Unable to initialize LLM ... LiteLLM fallback package is not installedCrewAI did not choose its OpenAI client and fell back to LiteLLM. Add custom_openai=True and base_url. Installing crewai[litellm] also silences it, but then LiteLLM does the routing.
The error mentions another provider's API keyThe model string matched one of CrewAI's own providers. Add custom_openai=True.
401 missing_api_key or invalid_api_keyThe key did not reach the request. Pass api_key= explicitly and check TOKENS_API_KEY is set in the process that runs the crew. See API keys.
404 model_not_foundThe id is wrong, or CrewAI stripped an openai/ that belonged to the id. Copy the id from /models; for ids that start with openai/, write the prefix twice.
402 insufficient_creditsThe plan credits and wallet cannot cover the request. Top up in billing. Loops that run up to max_iter steps use credits fast.
429 rate_limited or concurrency_limitToo many requests. Set max_rpm, run tasks in sequence, wait the Retry-After seconds. window_exhausted means a plan window is used up, so retrying will not help until it resets.
TimeoutsRaise timeout on the LLM, or stream. A 504 upstream_timeout comes from the upstream after 600 seconds. Try a smaller request.

Every code is in Errors. When you ask support for help, include the x-tokens-request-id response header. The examples on this page do not expose response headers, so reproduce the call once with cURL and -i to read it.

Was this page helpful?

Still stuck? Open a support ticket

Need help configuring your agent?

Test your connection with the connection tester, or create an API key.