Tool calling (function calling) মানে model আপনার code-কে বলে একটা function চালাতে, তারপর সেই ফলাফল কাজে লাগায়। chat completions আর messages-এ tool definition আর tool call আপনার আর provider-এর মাঝে Tokens পৌঁছে দেয়। তাই OpenAI বা Anthropic-এর জন্য যে code লিখতেন, এখানেও সেটাই চলবে। তবে কতটা ভালো চলবে, সেটা model-এর ওপর নির্ভর করে।
Tool calling কীভাবে কাজ করে#
- আপনি কথোপকথনের সাথে
tools-এর একটা তালিকা পাঠান। প্রতিটা tool-এর একটা নাম, বর্ণনা আর argument-এর জন্য JSON Schema থাকে। - model হয় text-এ উত্তর দেয়, নয়তো argument সহ এক বা একাধিক tool call ফেরত দেয়।
- আপনার code প্রতিটা function চালায় আর ফলাফল
toolmessage হিসেবে ফেরত পাঠায়। - model কোনো tool call না করে উত্তর দেওয়া পর্যন্ত এটা চলতে থাকে।
gateway কখনো tool চালায় না। সে শুধু message এদিক-ওদিক পৌঁছে দেয়।
/v1/chat/completions-এ OpenAI-style tool#
একটা tool definition এরকম:
{
"type": "function",
"function": {
"name": "get_current_time",
"description": "Get the current time in an IANA timezone, e.g. Asia/Dhaka.",
"parameters": {
"type": "object",
"properties": {
"timezone": { "type": "string", "description": "IANA timezone name" }
},
"required": ["timezone"]
}
}
}model এটা call করলে assistant message-এ content-এর বদলে (বা তার সাথে) tool_calls থাকে, আর finish_reason হয় "tool_calls":
{
"role": "assistant",
"content": null,
"tool_calls": [
{
"id": "call_01",
"type": "function",
"function": { "name": "get_current_time", "arguments": "{\"timezone\": \"Asia/Dhaka\"}" }
}
]
}arguments একটা JSON string, object নয়। তাই parse করে নিন। আর মনে রাখবেন, model মাঝে মাঝে ভুল JSON-ও বানিয়ে ফেলতে পারে।
Python-এ পুরো tool loop#
pip install openai করে আর TOKENS_API_KEY set করলে এটা যেমন আছে তেমনই চলবে। Windows-এ timezone data-র জন্য pip install tzdata-ও করে নিন। tool-গুলো শুধু standard library ব্যবহার করে, তাই আর কিছু configure করতে হবে না।
import json
import os
from datetime import datetime
from zoneinfo import ZoneInfo
from openai import OpenAI
client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"])
MODEL = "deepseek/deepseek-v4.1-flash"
def get_current_time(timezone: str) -> dict:
return {"timezone": timezone, "time": datetime.now(ZoneInfo(timezone)).isoformat()}
def add(a: float, b: float) -> dict:
return {"result": a + b}
FUNCTIONS = {"get_current_time": get_current_time, "add": add}
TOOLS = [
{
"type": "function",
"function": {
"name": "get_current_time",
"description": "Get the current time in an IANA timezone, e.g. Asia/Dhaka.",
"parameters": {
"type": "object",
"properties": {"timezone": {"type": "string"}},
"required": ["timezone"],
},
},
},
{
"type": "function",
"function": {
"name": "add",
"description": "Add two numbers.",
"parameters": {
"type": "object",
"properties": {"a": {"type": "number"}, "b": {"type": "number"}},
"required": ["a", "b"],
},
},
},
]
messages = [
{"role": "user", "content": "What time is it in Dhaka and in London? Also, what is 1250.5 + 349.5?"}
]
for _ in range(8): # hard stop so a confused model can't loop forever
resp = client.chat.completions.create(
model=MODEL, messages=messages, tools=TOOLS, tool_choice="auto", max_tokens=1024
)
msg = resp.choices[0].message
messages.append(msg.model_dump(exclude_none=True))
if not msg.tool_calls:
print(msg.content)
break
for call in msg.tool_calls:
fn = FUNCTIONS.get(call.function.name)
try:
args = json.loads(call.function.arguments or "{}")
result = fn(**args) if fn else {"error": f"unknown tool {call.function.name}"}
except Exception as exc: # report errors back to the model instead of crashing
result = {"error": str(exc)}
messages.append(
{"role": "tool", "tool_call_id": call.id, "content": json.dumps(result)}
)
else:
print("Stopped after 8 rounds without a final answer.")কয়েকটা ছোট জিনিস মাথায় রাখলে debug করার সময় বাঁচবে:
tool_calls-ওয়ালা assistant message-টা আগে append করুন, তারপরtoolresult। যেtoolmessage-এরtool_call_idআগের কোনো call-এর সাথে মেলে না, বেশিরভাগ provider সেটা reject করে।- model এক turn-এ একাধিক tool call ফেরত দিতে পারে। পরের request-এর আগে সবগুলোর উত্তর দিন।
- error হলে সেটা tool result হিসেবেই model-কে ফেরত দিন। অনেক সময় সে নিজেই সামলে নেয়, যেমন ভুল timezone নাম ঠিক করে আবার চেষ্টা করে।
- প্রতিটা round আলাদা billed request, আর প্রতিবার পুরো কথোপকথন আবার পাঠাতে হয়। তাই loop লম্বা হলে খরচ শেষ উত্তর দেখে যা মনে হয়, তার চেয়ে বেশি পড়ে।
tool_choice দিয়ে tool ব্যবহার নিয়ন্ত্রণ#
| Value | Effect |
|---|---|
"auto" | model নিজে ঠিক করে (tool থাকলে এটাই default) |
"none" | model-কে text-এই উত্তর দিতে হবে |
"required" | model-কে অন্তত একটা tool call করতেই হবে |
{"type": "function", "function": {"name": "add"}} | model-কে ঠিক ওই function-টাই call করতে হবে |
"required" আর নির্দিষ্ট function জোর করে call করানো সব model আর provider-এ একরকম চলে না। কেউ কেউ এটা এড়িয়ে যায়, কেউ কেউ 400 ফেরত দেয়। জোর করে call করানো লাগলে আগে ঠিক যে model ব্যবহার করবেন সেটাতেই পরীক্ষা করে নিন।
/v1/messages-এ Anthropic-style tool#
/v1/messages-এ Anthropic-এর format-ই চলবে। tool-এ থাকে name, description আর input_schema। model উত্তর দেয় tool_use content block দিয়ে, আর আপনি user message-এ tool_result block দিয়ে জবাব দেন।
import os
import anthropic
client = anthropic.Anthropic(base_url="https://tokens.bd", api_key=os.environ["TOKENS_API_KEY"])
tools = [{
"name": "add",
"description": "Add two numbers.",
"input_schema": {
"type": "object",
"properties": {"a": {"type": "number"}, "b": {"type": "number"}},
"required": ["a", "b"],
},
}]
messages = [{"role": "user", "content": "What is 1250.5 + 349.5?"}]
resp = client.messages.create(
model="deepseek/deepseek-v4.1-flash", max_tokens=1024, tools=tools, messages=messages
)
if resp.stop_reason == "tool_use":
call = next(b for b in resp.content if b.type == "tool_use")
total = call.input["a"] + call.input["b"]
messages += [
{"role": "assistant", "content": resp.content},
{"role": "user", "content": [
{"type": "tool_result", "tool_use_id": call.id, "content": str(total)}
]},
]
resp = client.messages.create(
model="deepseek/deepseek-v4.1-flash", max_tokens=1024, tools=tools, messages=messages
)
print(resp.content[0].text)এখানে tool_choice Anthropic-এর ধরনেই দিতে হয়: {"type": "auto"}, {"type": "any"} বা {"type": "tool", "name": "add"}। Claude নয় এমন কোনো model /v1/messages-এ চালালে gateway tool আর tool call OpenAI format-এ আর সেখান থেকে ফেরত translate করে। translate হওয়ার পর কী কী টিকে থাকে, তা messages পেজে দেখুন। anthropic-beta header সেই provider-দের কাছে forward হয় যারা Messages API নিজেরাই বোঝে। তাই যে server-side tool-এ এই header লাগে, সেটা সেই provider-এর support থাকলে চলবে।
Tool calling-এর টিপ#
- সব model-এ tool চলে না। কোনো model-এর ওপর agent বানানোর আগে catalog-এ সেই model-এর পেজ দেখে নিন। ছোট বা পুরনো model tool এড়িয়ে যেতে পারে, নয়তো argument বানিয়ে ফেলতে পারে।
- schema সহজ রাখুন। গভীরভাবে nested schema-র চেয়ে সমতল object আর পরিষ্কার বর্ণনা দিলে argument বেশি নির্ভরযোগ্য আসে।
- Streaming-এও চলে। tool call-এর argument ছোট ছোট অংশে আসে
delta.tool_calls-এ (messages-এinput_json_delta-তে)। parse করার আগে অংশগুলো জোড়া লাগিয়ে নিন। streaming পেজ দেখুন। - চালানোর আগে validate করুন। argument-কে অবিশ্বাস্য input ধরে নিন, বিশেষ করে যে tool file, shell বা টাকাপয়সা ছোঁয়।
একটা model-এ tool call 400 দিয়ে fail করছে অথচ আরেকটায় চলছে, মানে ওই model বা তার provider সেই tool feature নেয় না। আরও সাধারণ সমস্যার কথা troubleshooting পেজে আছে।