Skip to content

Tool Calling

Use OpenAI-style function tools and Anthropic-style tools through the Tokens gateway, with a complete runnable tool loop in Python and tips on model support.

On this page

Tool calling (function calling) lets a model ask your code to run a function and then use the result. Tokens carries tool definitions and tool calls between you and the provider on chat completions and messages, so the same code you would write against OpenAI or Anthropic works here. Whether it works well depends on the model.

How tool calling works#

  1. You send the conversation plus a list of tools, each with a name, description and JSON Schema for its arguments.
  2. The model either answers in text or returns one or more tool calls with arguments.
  3. Your code runs each function and sends the results back as tool messages.
  4. Repeat until the model answers without calling a tool.

The gateway never executes tools. It only carries the messages.

OpenAI-style tools on /v1/chat/completions#

A tool definition:

json
{
  "type": "function",
  "function": {
    "name": "get_current_time",
    "description": "Get the current time in an IANA timezone, e.g. Asia/Dhaka.",
    "parameters": {
      "type": "object",
      "properties": {
        "timezone": { "type": "string", "description": "IANA timezone name" }
      },
      "required": ["timezone"]
    }
  }
}

When the model calls it, the assistant message has tool_calls instead of (or alongside) content, and finish_reason is "tool_calls":

json
{
  "role": "assistant",
  "content": null,
  "tool_calls": [
    {
      "id": "call_01",
      "type": "function",
      "function": { "name": "get_current_time", "arguments": "{\"timezone\": \"Asia/Dhaka\"}" }
    }
  ]
}

arguments is a JSON string, not an object. Parse it, and expect that a model can occasionally produce invalid JSON.

Complete tool loop in Python#

This runs as is with pip install openai and TOKENS_API_KEY set (on Windows, also pip install tzdata for timezone data). The tools use only the standard library, so there is nothing else to configure.

tool_loop.py
import json
import os
from datetime import datetime
from zoneinfo import ZoneInfo

from openai import OpenAI

client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"])
MODEL = "deepseek/deepseek-v4.1-flash"


def get_current_time(timezone: str) -> dict:
    return {"timezone": timezone, "time": datetime.now(ZoneInfo(timezone)).isoformat()}


def add(a: float, b: float) -> dict:
    return {"result": a + b}


FUNCTIONS = {"get_current_time": get_current_time, "add": add}

TOOLS = [
    {
        "type": "function",
        "function": {
            "name": "get_current_time",
            "description": "Get the current time in an IANA timezone, e.g. Asia/Dhaka.",
            "parameters": {
                "type": "object",
                "properties": {"timezone": {"type": "string"}},
                "required": ["timezone"],
            },
        },
    },
    {
        "type": "function",
        "function": {
            "name": "add",
            "description": "Add two numbers.",
            "parameters": {
                "type": "object",
                "properties": {"a": {"type": "number"}, "b": {"type": "number"}},
                "required": ["a", "b"],
            },
        },
    },
]

messages = [
    {"role": "user", "content": "What time is it in Dhaka and in London? Also, what is 1250.5 + 349.5?"}
]

for _ in range(8):  # hard stop so a confused model can't loop forever
    resp = client.chat.completions.create(
        model=MODEL, messages=messages, tools=TOOLS, tool_choice="auto", max_tokens=1024
    )
    msg = resp.choices[0].message
    messages.append(msg.model_dump(exclude_none=True))

    if not msg.tool_calls:
        print(msg.content)
        break

    for call in msg.tool_calls:
        fn = FUNCTIONS.get(call.function.name)
        try:
            args = json.loads(call.function.arguments or "{}")
            result = fn(**args) if fn else {"error": f"unknown tool {call.function.name}"}
        except Exception as exc:  # report errors back to the model instead of crashing
            result = {"error": str(exc)}
        messages.append(
            {"role": "tool", "tool_call_id": call.id, "content": json.dumps(result)}
        )
else:
    print("Stopped after 8 rounds without a final answer.")

A few details that save debugging time:

  • Append the assistant message that contains tool_calls before the tool results. Most providers reject a tool message whose tool_call_id doesn't match a preceding call.
  • Models can return several tool calls in one turn. Answer all of them before the next request.
  • Errors go back to the model as tool results. It can often recover, for example by fixing a bad timezone name.
  • Every round is a separate billed request that resends the whole conversation, so long loops cost more than the final answer suggests.

Control tool use with tool_choice#

ValueEffect
"auto"Model decides (the default when tools are present)
"none"Model must answer in text
"required"Model must call at least one tool
{"type": "function", "function": {"name": "add"}}Model must call that function

Support for "required" and forced functions varies by model and provider. Some ignore it; some return a 400. If you need a forced call, test it on the exact model first.

Anthropic-style tools on /v1/messages#

On /v1/messages, use Anthropic's format: tools have name, description and input_schema, the model replies with tool_use content blocks, and you answer with tool_result blocks in a user message.

python
import os
import anthropic

client = anthropic.Anthropic(base_url="https://tokens.bd", api_key=os.environ["TOKENS_API_KEY"])

tools = [{
    "name": "add",
    "description": "Add two numbers.",
    "input_schema": {
        "type": "object",
        "properties": {"a": {"type": "number"}, "b": {"type": "number"}},
        "required": ["a", "b"],
    },
}]
messages = [{"role": "user", "content": "What is 1250.5 + 349.5?"}]

resp = client.messages.create(
    model="deepseek/deepseek-v4.1-flash", max_tokens=1024, tools=tools, messages=messages
)
if resp.stop_reason == "tool_use":
    call = next(b for b in resp.content if b.type == "tool_use")
    total = call.input["a"] + call.input["b"]
    messages += [
        {"role": "assistant", "content": resp.content},
        {"role": "user", "content": [
            {"type": "tool_result", "tool_use_id": call.id, "content": str(total)}
        ]},
    ]
    resp = client.messages.create(
        model="deepseek/deepseek-v4.1-flash", max_tokens=1024, tools=tools, messages=messages
    )
print(resp.content[0].text)

tool_choice takes Anthropic's forms here: {"type": "auto"}, {"type": "any"} or {"type": "tool", "name": "add"}. When a non-Claude model is served through /v1/messages, the gateway translates tools and tool calls to and from the OpenAI format; see messages for what survives translation. Server-side tools that need the anthropic-beta header won't be enabled, because that header isn't forwarded.

Tips for tool calling#

  • Not every model supports tools. Check the model's page in the catalog before building an agent on it. Small or older models may ignore tools or invent arguments.
  • Keep schemas simple. Flat objects with clear descriptions get more reliable arguments than deeply nested schemas.
  • Streaming works. Tool call arguments arrive as fragments in delta.tool_calls (or input_json_delta on messages); concatenate them before parsing. See streaming.
  • Validate before executing. Treat arguments as untrusted input, especially for tools that touch files, shells or money.

If tool calls fail with a 400 on one model but work on another, the model or its provider doesn't accept that tool feature. The troubleshooting page covers other common failures.

Was this page helpful?

Still stuck? Open a support ticket

Need help configuring your agent?

Test your connection with the connection tester, or create an API key.