Structured output মানে model-কে গদ্যের বদলে এমন JSON দিতে বলা, যা আপনার code সরাসরি parse করতে পারে। Tokens দিয়ে এটা করার চারটা পথ আছে। কোনটা চলবে, সেটা নির্ভর করে model-এর ওপর আর gateway কীভাবে সেই model-এ পৌঁছায় তার ওপর:
| Method | Endpoint | Guarantee কতটা শক্ত |
|---|---|---|
JSON mode: response_format: {"type": "json_object"} | Chat Completions | Valid JSON পাবেন, কিন্তু কোন key থাকবে তার কোনো কথা নেই |
JSON Schema: response_format: {"type": "json_schema", ...} | Chat Completions | Output আপনার schema মানে (strict: true দিলে), যদি model support করে |
| A forced tool call whose arguments are your schema | Chat Completions, Messages | Arguments schema মানে, model যতটা পারে; বেশি model-এ চলে |
output_config.format with a JSON Schema | Messages | Output আপনার schema মানে, যেসব model support করে তাদের ক্ষেত্রে |
এর কোনোটাই Tokens নিজে enforce করে না। gateway JSON পড়ে না, validate করে না, মেরামতও করে না। request-এর field forward করে আর provider-এর উত্তর ফেরত দেয়। তাই উপরের guarantee-গুলো provider-এর, Tokens-এর নয়। যে model কোনো feature support করে না, সে হয় field-টা ignore করে, নয়তো request ফিরিয়ে দেয়। উত্তর পেলে নিজের code-এ সবসময় parse আর validate করে নিন।
Gateway কী করে আর কী বাদ দেয়#
- OpenAI format বোঝে এমন provider-এ Chat Completions গেলে
response_formatঅপরিবর্তিত forward হয়। - Anthropic format বোঝে এমন provider-এ Messages গেলে
output_config,toolsআরtool_choiceঅপরিবর্তিত forward হয়। - কোনো model শুধু অন্য format-এর provider-এর কাছে পাওয়া গেলে gateway request translate করে, আর translation-এ যায় শুধু একটা নির্দিষ্ট তালিকার field। Anthropic-format provider-এ Chat Completions request গেলে
response_formatযায় না (n,seed,logprobsআর penalty-র field-গুলোও যায় না)। OpenAI-format provider-এ Messages request গেলেoutput_configযায় না। দুই ক্ষেত্রেই field-টার কোনো কাজ হয় না, আর কেউ আপনাকে জানায়ও না। model শুধু সাধারণ text-এ উত্তর দেয়। - Tools দুই দিকের translation-ই টিকে যায়, আর সেজন্যই forced tool call সবচেয়ে বেশি জায়গায় চলে। Tool calling আর Messages দেখুন।
structured request-এর উত্তরে গদ্য এলে সম্ভাব্য কারণ দুটো: model support করে না, অথবা translate হওয়া পথে field বাদ পড়ে গেছে। একই model-এ tool-call পদ্ধতিটা চেষ্টা করে দেখুন।
Model support করে কি না কীভাবে জানবেন#
/models পেজের model catalog-এ context window, দাম আর বর্ণনা আছে, কিন্তু JSON mode বা JSON Schema support-এর জন্য কোনো field নেই। যাচাই করার উপায়:
- model-এর maker-এর documentation পড়ুন (যেমন সেখানে "JSON output" বা "structured outputs" লেখা আছে কি না)।
- field-সহ একটা ছোট request পাঠিয়ে ফল দেখুন:
400এলে model বা provider সেটা নেয় না, মুক্ত গদ্য এলে field ignore হয়েছে, valid JSON এলে সেই request-এ কাজ করছে। - test কয়েকবার চালান। শর্তহীন JSON mode একবার পাস করে পরের prompt-এ fail করতে পারে।
Tokens যেসব model সবচেয়ে বেশি তুলনা করে, সেগুলোর তালিকা Model বেছে নেওয়া পেজে আছে।
JSON mode#
JSON mode-এ model একটা valid JSON object ফেরত দেয়, কিন্তু তাতে কোন key থাকবে তার কোনো নিশ্চয়তা নেই। তাই কোনো একটা message-এ model-কে স্পষ্ট বলে দিন JSON দিতে, আর কেমন shape চান তা বর্ণনা করুন। OpenAI-এর documentation বলছে, conversation-এর কোনো একটা message-এ JSON বানাতে বলতেই হবে। না বললে model token limit না আসা পর্যন্ত শুধু whitespace দিয়ে উত্তর ভরে ফেলতে পারে, আর context-এ "JSON" শব্দটা না থাকলে API request ফিরিয়েও দিতে পারে।
curl https://tokens.bd/v1/chat/completions \
-H "Authorization: Bearer $TOKENS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek/deepseek-v4.1-flash",
"max_tokens": 400,
"response_format": {"type": "json_object"},
"messages": [
{"role": "system", "content": "Reply with a JSON object with the keys \"language\" and \"summary\"."},
{"role": "user", "content": "def add(a, b): return a + b"}
]
}'import json
import os
from openai import OpenAI
client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"])
resp = client.chat.completions.create(
model="deepseek/deepseek-v4.1-flash",
max_tokens=400,
response_format={"type": "json_object"},
messages=[
{"role": "system", "content": 'Reply with a JSON object with the keys "language" and "summary".'},
{"role": "user", "content": "def add(a, b): return a + b"},
],
)
choice = resp.choices[0]
if choice.finish_reason == "length":
raise RuntimeError("Cut off by max_tokens: the JSON is incomplete")
data = json.loads(choice.message.content)
print(data)import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://tokens.bd/v1", apiKey: process.env.TOKENS_API_KEY });
const resp = await client.chat.completions.create({
model: "deepseek/deepseek-v4.1-flash",
max_tokens: 400,
response_format: { type: "json_object" },
messages: [
{ role: "system", content: 'Reply with a JSON object with the keys "language" and "summary".' },
{ role: "user", content: "def add(a, b): return a + b" },
],
});
const choice = resp.choices[0];
if (choice.finish_reason === "length") throw new Error("Cut off by max_tokens: the JSON is incomplete");
console.log(JSON.parse(choice.message.content ?? "{}"));parse করার আগে finish_reason দেখে নিন। সেটা length হলে বুঝবেন model object-এর মাঝখানে max_tokens-এ ঠেকে গেছে, তাই JSON কাটা পড়েছে। reasoning model ওই limit-এর একটা অংশ আগে ভাবতেই খরচ করে, তাই তাদের বেশি জায়গা দিন। Reasoning ও thinking model পেজ দেখুন।
response_format-এ JSON Schema#
JSON Schema output-এ উত্তর আপনার দেওয়া schema-র ভেতরে বাঁধা থাকে। Chat Completions-এ request দেখতে এরকম (shape-টা OpenAI-র, October 2026-এ দেখা হয়েছে OpenAI-এর structured outputs guide আর migration guide থেকে):
{
"response_format": {
"type": "json_schema",
"json_schema": {
"name": "ticket",
"strict": true,
"schema": {
"type": "object",
"properties": {
"title": { "type": "string" },
"severity": { "type": "string", "enum": ["low", "medium", "high"] },
"needs_followup": { "type": "boolean" }
},
"required": ["title", "severity", "needs_followup"],
"additionalProperties": false
}
}
}
}strict: true দিলে OpenAI schema-র ওপর এই নিয়মগুলো খাটায়, আর নিয়ম ভাঙলে schema ফিরিয়ে দেয়:
- Root অবশ্যই object হতে হবে, আর সেটা
anyOfহতে পারবে না। - প্রতিটা property
required-এ থাকতে হবে। কোনো field optional করতে চাইলেnullমঞ্জুর করুন:{"type": ["string", "null"]}। - প্রতিটা object-এ
"additionalProperties": falseলাগবে। - সব মিলিয়ে সর্বোচ্চ 5,000টা property আর 10 স্তর nesting।
অন্য maker-রা JSON Schema-র একটা অংশ support করে, আর সেই অংশটা একেক জায়গায় একেক রকম। provider আপনার schema ফিরিয়ে দিলে সেটা সহজ করুন: সমতল object, সাধারণ type, আর নির্দিষ্ট choice-এর জন্য enum।
import json
import os
from openai import OpenAI
client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"])
schema = {
"type": "object",
"properties": {
"title": {"type": "string"},
"severity": {"type": "string", "enum": ["low", "medium", "high"]},
"needs_followup": {"type": "boolean"},
},
"required": ["title", "severity", "needs_followup"],
"additionalProperties": False,
}
resp = client.chat.completions.create(
model="deepseek/deepseek-v4.1-flash",
max_tokens=400,
messages=[
{"role": "user", "content": "Users report that the login page returns a 500 after the last deploy."}
],
response_format={
"type": "json_schema",
"json_schema": {"name": "ticket", "strict": True, "schema": schema},
},
)
choice = resp.choices[0]
if getattr(choice.message, "refusal", None):
print("The model refused:", choice.message.refusal)
elif choice.finish_reason != "stop":
print("Incomplete answer, finish_reason =", choice.finish_reason)
else:
ticket = json.loads(choice.message.content)
print(ticket["severity"], ticket["title"])import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://tokens.bd/v1", apiKey: process.env.TOKENS_API_KEY });
const schema = {
type: "object",
properties: {
title: { type: "string" },
severity: { type: "string", enum: ["low", "medium", "high"] },
needs_followup: { type: "boolean" },
},
required: ["title", "severity", "needs_followup"],
additionalProperties: false,
};
const resp = await client.chat.completions.create({
model: "deepseek/deepseek-v4.1-flash",
max_tokens: 400,
messages: [
{ role: "user", content: "Users report that the login page returns a 500 after the last deploy." },
],
response_format: { type: "json_schema", json_schema: { name: "ticket", strict: true, schema } },
});
const choice = resp.choices[0];
if (choice.finish_reason !== "stop") {
console.log("Incomplete or refused answer:", choice.finish_reason, choice.message.refusal);
} else {
console.log(JSON.parse(choice.message.content ?? "{}"));
}curl https://tokens.bd/v1/chat/completions \
-H "Authorization: Bearer $TOKENS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek/deepseek-v4.1-flash",
"max_tokens": 400,
"messages": [
{"role": "user", "content": "Users report that the login page returns a 500 after the last deploy."}
],
"response_format": {
"type": "json_schema",
"json_schema": {
"name": "ticket",
"strict": true,
"schema": {
"type": "object",
"properties": {
"title": {"type": "string"},
"severity": {"type": "string", "enum": ["low", "medium", "high"]},
"needs_followup": {"type": "boolean"}
},
"required": ["title", "severity", "needs_followup"],
"additionalProperties": false
}
}
}
}'parse করার আগে তিনটা ফলাফল সামলে নিন: refusal (OpenAI এটা message.refusal-এ জানায), অসম্পূর্ণ উত্তর (finish_reason হলো length বা content_filter), আর এমন উত্তর যা parse হয় কিন্তু আপনার নিজের check পাস করে না। OpenAI আর Anthropic-এর SDK-তে এমন helper আছে যা Pydantic model বা Zod schema থেকে schema বানিয়ে দেয় আর উত্তর parse করে। তারা উপরের মতো একই JSON পাঠায়, তাই model feature-টা support করলে Tokens দিয়েও চলে।
Responses API
Responses API একই schema নেয় response_format-এর বদলে text.format-এর নিচে, যেখানে name, strict আর schema এক ধাপ ওপরে থাকে, ভেতরে আলাদা json_schema key থাকে না। Responses API দেখুন। Support নির্ভর করে model-এর পেছনের provider-এর ওপর।
Tools দিয়ে structured output#
forced tool call সবচেয়ে বেশি জায়গায় চলে, কারণ response_format-এর চেয়ে অনেক বেশি model tool calling support করে, আর এটা gateway-র format translation-এও টিকে যায়। একটা function বানান যার parameters আপনার চাওয়া schema, model-কে সেটা call করতে বাধ্য করুন, তারপর arguments পড়ে নিন।
import json
import os
from openai import OpenAI
client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"])
tools = [
{
"type": "function",
"function": {
"name": "record_ticket",
"description": "Record a support ticket extracted from the user's message.",
"strict": True,
"parameters": {
"type": "object",
"properties": {
"title": {"type": "string"},
"severity": {"type": "string", "enum": ["low", "medium", "high"]},
},
"required": ["title", "severity"],
"additionalProperties": False,
},
},
}
]
resp = client.chat.completions.create(
model="deepseek/deepseek-v4.1-flash",
max_tokens=400,
messages=[{"role": "user", "content": "The login page returns a 500 after the last deploy."}],
tools=tools,
tool_choice={"type": "function", "function": {"name": "record_ticket"}},
)
call = resp.choices[0].message.tool_calls[0]
ticket = json.loads(call.function.arguments)
print(ticket)Chat Completions-এ arguments একটা JSON string, তাই এটা parse করতে হবে। function object-এর ভেতরের strict হলো schema হুবহু মানানোর জন্য OpenAI-র switch, আর এর schema-র নিয়ম উপরের মতোই। যেসব model ও provider এটা চেনে না, তারা ignore করতে পারে। এখানে কিছুই Tokens চালায় না, Tokens শুধু call-টা বয়ে নিয়ে যায়। নির্দিষ্ট একটা function-এ tool_choice বেঁধে দেওয়ার support model ভেদে আলাদা: কেউ ignore করে, কেউ 400 দেয়। Tool calling দেখুন।
OpenAI-র documentation বলছে, GPT-5.4 আর তার পরের model-এ reasoning_effort none ছাড়া অন্য কিছু হলে Chat Completions tool calling support করে না। কোনো OpenAI reasoning model-এ forced tool call fail করলে reasoning_effort none করে দিন, অথবা Responses API ব্যবহার করুন।
Messages-এ একই ধারণা চলে Anthropic-এর tool format-এ, আর arguments tool_use block-এ আগে থেকেই object হয়ে parse করা অবস্থায় ফেরত আসে:
import os
import anthropic
client = anthropic.Anthropic(base_url="https://tokens.bd", api_key=os.environ["TOKENS_API_KEY"])
tools = [
{
"name": "record_ticket",
"description": "Record a support ticket extracted from the user's message.",
"input_schema": {
"type": "object",
"properties": {
"title": {"type": "string"},
"severity": {"type": "string", "enum": ["low", "medium", "high"]},
},
"required": ["title", "severity"],
},
}
]
message = client.messages.create(
model="deepseek/deepseek-v4.1-flash",
max_tokens=400,
tools=tools,
tool_choice={"type": "tool", "name": "record_ticket"},
messages=[{"role": "user", "content": "The login page returns a 500 after the last deploy."}],
)
block = next(b for b in message.content if b.type == "tool_use")
print(block.input)Anthropic জানায়, tool_choice-এ any বা tool দিয়ে tool জোর করলে তাকে তাদের manual extended thinking (thinking.type: "enabled")-এর সাথে একসাথে চালানো যায় না, আর তাদের একেবারে নতুন কয়েকটা model adaptive thinking থাকলেও forced tool use নেয় না। Claude model-এ thinking আর structured output দুটোই চাইলে আগে Reasoning ও thinking model পড়ুন।
Messages-এ output_config দিয়ে JSON Schema#
Anthropic-এর Messages API-র নিজস্ব JSON Schema output আছে, যা output_config.format-এ দিতে হয় (October 2026-এ দেখা হয়েছে Anthropic-এর structured outputs guide থেকে)। উত্তর আসে text block-এর ভেতর text হিসেবে, সেই text-টাই JSON। Anthropic tool definition-এ strict: true-ও দেয়, যাতে tool-এর নাম আর input schema-র সাথে মিলে যাওয়া নিশ্চিত হয়।
curl https://tokens.bd/v1/messages \
-H "x-api-key: $TOKENS_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "deepseek/deepseek-v4.1-flash",
"max_tokens": 400,
"messages": [
{"role": "user", "content": "The login page returns a 500 after the last deploy."}
],
"output_config": {
"format": {
"type": "json_schema",
"schema": {
"type": "object",
"properties": {
"title": {"type": "string"},
"severity": {"type": "string", "enum": ["low", "medium", "high"]}
},
"required": ["title", "severity"],
"additionalProperties": false
}
}
}
}'import json
import os
import anthropic
client = anthropic.Anthropic(base_url="https://tokens.bd", api_key=os.environ["TOKENS_API_KEY"])
response = client.messages.create(
model="deepseek/deepseek-v4.1-flash",
max_tokens=400,
messages=[{"role": "user", "content": "The login page returns a 500 after the last deploy."}],
output_config={
"format": {
"type": "json_schema",
"schema": {
"type": "object",
"properties": {
"title": {"type": "string"},
"severity": {"type": "string", "enum": ["low", "medium", "high"]},
},
"required": ["title", "severity"],
"additionalProperties": False,
},
}
},
)
if response.stop_reason in ("refusal", "max_tokens"):
raise RuntimeError(f"No valid JSON: stop_reason = {response.stop_reason}")
text = next(b.text for b in response.content if b.type == "text")
print(json.loads(text))এর জন্য Anthropic SDK-র নতুন version লাগে, যে version output_config চেনে। আপনার SDK argument-টা ফিরিয়ে দিলে SDK upgrade করুন। Anthropic কোন কোন Claude model-এ feature-টা চলে তার তালিকা দেয়, আর JSON Schema-র শুধু একটা অংশ support করে: recursive schema চলে না, সংখ্যা বা string-এর দৈর্ঘ্যের constraint চলে না, object-এ additionalProperties: false দিতে হয়, আর unsupported feature দিলে 400 আসে। stop_reason refusal হলেও 200 আসে আর bill হয়, তখন output schema না-ও মানতে পারে; max_tokens-এ ঠেকলে উত্তর অসম্পূর্ণ থাকতে পারে।
Tokens-এ এই পথের সীমাটা মনে রাখুন। model-টা যে provider-এর কাছে পাওয়া যায় সে Anthropic format বুঝলে এটা কাজ করে। gateway আপনার Messages request OpenAI format-এ translate করলে output_config বাদ পড়ে যায় (উপরে দেখুন)। সব model-এর জন্য একটাই code path চাইলে forced-tool পদ্ধতি ব্যবহার করুন।
Streaming আর structured output#
"stream": true দিলেও structured output চলে, কিন্তু JSON valid হয় stream পুরো শেষ হলে। content delta-গুলো (বা tool-call arguments-এর টুকরো, বা Anthropic-এর input_json_delta অংশগুলো) জোড়া দিন, আর শেষে একবার parse করুন। অর্ধেক text parse করতে যাবেন না। streaming দেখুন।
খরচ আর limit#
- schema, tool definition আর আপনার instruction প্রতিটা request-এ input token হিসেবে গোনা হয়। বড় schema প্রতিটা call-এ খরচ বাড়ায়।
- Tokens provider-এর report করা usage-ই bill করে, model-এর input আর output price ধরে। structured output-এর জন্য বাড়তি কোনো চার্জ নেই।
max_tokensপুরো উত্তরের সীমা। কাটা পড়া JSON object parse হয় না, তাই আপনার সবচেয়ে লম্বা সম্ভাব্য উত্তরের চেয়ে বেশি রাখুন, আর model আগে ভাবলে আরও জায়গা যোগ করুন।- request forward করার আগে gateway
max_tokensধরে worst-case খরচ reserve করে। reservation নিয়ে Chat Completions পেজ দেখুন।
নির্ভরযোগ্য করার উপায়#
- strict mode ব্যবহার করলেও প্রতিটা উত্তর নিজের code-এ নিজের schema-র সাথে validate করুন (Pydantic, Zod,
jsonschema)। - schema সমতল রাখুন আর প্রতিটা field-এ ছোট একটা
descriptionদিন; model সেটা পড়ে। - extraction-এর কাজে
temperatureকম রাখুন। কিছু reasoning model এটা মানে না বা ignore করে; Reasoning ও thinking model দেখুন। - parse fail করলে error message জুড়ে একবার retry করুন, তারপর স্পষ্টভাবে fail করুন। প্রতিটা retry নতুন bill-হওয়া request।
- production-এ যে model চালাবেন, ঠিক সেটাতেই test করুন। এক model-এ চলা schema আরেক model ফিরিয়ে দিতে পারে।