# Structured output: JSON mode ও JSON Schema

> Tokens দিয়ে model থেকে machine-readable JSON পাওয়ার উপায়: JSON mode, response_format-এ JSON Schema, forced tool call আর Anthropic-এর output_config, সাথে উদাহরণ, validation-এর পরামর্শ আর gateway কী বাদ দেয়।

Structured output মানে model-কে গদ্যের বদলে এমন JSON দিতে বলা, যা আপনার code সরাসরি parse করতে পারে। Tokens দিয়ে এটা করার চারটা পথ আছে। কোনটা চলবে, সেটা নির্ভর করে model-এর ওপর আর gateway কীভাবে সেই model-এ পৌঁছায় তার ওপর:

| Method                                                       | Endpoint                   | Guarantee কতটা শক্ত                                                                         |
| ------------------------------------------------------------ | -------------------------- | ------------------------------------------------------------------------------------------- |
| JSON mode: `response_format: {"type": "json_object"}`        | Chat Completions           | Valid JSON পাবেন, কিন্তু কোন key থাকবে তার কোনো কথা নেই                                      |
| JSON Schema: `response_format: {"type": "json_schema", ...}` | Chat Completions           | Output আপনার schema মানে (`strict: true` দিলে), যদি model support করে                       |
| A forced tool call whose arguments are your schema           | Chat Completions, Messages | Arguments schema মানে, model যতটা পারে; বেশি model-এ চলে                                    |
| `output_config.format` with a JSON Schema                    | Messages                   | Output আপনার schema মানে, যেসব model support করে তাদের ক্ষেত্রে                              |

এর কোনোটাই Tokens নিজে enforce করে না। gateway JSON পড়ে না, validate করে না, মেরামতও করে না। request-এর field forward করে আর provider-এর উত্তর ফেরত দেয়। তাই উপরের guarantee-গুলো provider-এর, Tokens-এর নয়। যে model কোনো feature support করে না, সে হয় field-টা ignore করে, নয়তো request ফিরিয়ে দেয়। উত্তর পেলে নিজের code-এ সবসময় parse আর validate করে নিন।

## Gateway কী করে আর কী বাদ দেয়

- OpenAI format বোঝে এমন provider-এ **Chat Completions** গেলে `response_format` অপরিবর্তিত forward হয়।
- Anthropic format বোঝে এমন provider-এ **Messages** গেলে `output_config`, `tools` আর `tool_choice` অপরিবর্তিত forward হয়।
- কোনো model শুধু **অন্য** format-এর provider-এর কাছে পাওয়া গেলে gateway request translate করে, আর translation-এ যায় শুধু একটা নির্দিষ্ট তালিকার field। Anthropic-format provider-এ Chat Completions request গেলে `response_format` যায় না (`n`, `seed`, `logprobs` আর penalty-র field-গুলোও যায় না)। OpenAI-format provider-এ Messages request গেলে `output_config` যায় না। দুই ক্ষেত্রেই field-টার কোনো কাজ হয় না, আর কেউ আপনাকে জানায়ও না। model শুধু সাধারণ text-এ উত্তর দেয়।
- Tools দুই দিকের translation-ই টিকে যায়, আর সেজন্যই forced tool call সবচেয়ে বেশি জায়গায় চলে। [Tool calling](/docs/tool-calling) আর [Messages](/docs/messages) দেখুন।

structured request-এর উত্তরে গদ্য এলে সম্ভাব্য কারণ দুটো: model support করে না, অথবা translate হওয়া পথে field বাদ পড়ে গেছে। একই model-এ tool-call পদ্ধতিটা চেষ্টা করে দেখুন।

## Model support করে কি না কীভাবে জানবেন

[/models](/models) পেজের model catalog-এ context window, দাম আর বর্ণনা আছে, কিন্তু JSON mode বা JSON Schema support-এর জন্য কোনো field নেই। যাচাই করার উপায়:

1. model-এর maker-এর documentation পড়ুন (যেমন সেখানে "JSON output" বা "structured outputs" লেখা আছে কি না)।
2. field-সহ একটা ছোট request পাঠিয়ে ফল দেখুন: `400` এলে model বা provider সেটা নেয় না, মুক্ত গদ্য এলে field ignore হয়েছে, valid JSON এলে সেই request-এ কাজ করছে।
3. test কয়েকবার চালান। শর্তহীন JSON mode একবার পাস করে পরের prompt-এ fail করতে পারে।

Tokens যেসব model সবচেয়ে বেশি তুলনা করে, সেগুলোর তালিকা [Model বেছে নেওয়া](/docs/choosing-a-model) পেজে আছে।

## JSON mode

JSON mode-এ model একটা valid JSON object ফেরত দেয়, কিন্তু তাতে কোন key থাকবে তার কোনো নিশ্চয়তা নেই। তাই কোনো একটা message-এ model-কে স্পষ্ট বলে দিন JSON দিতে, আর কেমন shape চান তা বর্ণনা করুন। OpenAI-এর documentation বলছে, conversation-এর কোনো একটা message-এ JSON বানাতে বলতেই হবে। না বললে model token limit না আসা পর্যন্ত শুধু whitespace দিয়ে উত্তর ভরে ফেলতে পারে, আর context-এ "JSON" শব্দটা না থাকলে API request ফিরিয়েও দিতে পারে।

:::code-tabs

```bash title="cURL"
curl https://tokens.bd/v1/chat/completions \
  -H "Authorization: Bearer $TOKENS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek/deepseek-v4.1-flash",
    "max_tokens": 400,
    "response_format": {"type": "json_object"},
    "messages": [
      {"role": "system", "content": "Reply with a JSON object with the keys \"language\" and \"summary\"."},
      {"role": "user", "content": "def add(a, b): return a + b"}
    ]
  }'
```

```python title="Python"
import json
import os
from openai import OpenAI

client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"])

resp = client.chat.completions.create(
    model="deepseek/deepseek-v4.1-flash",
    max_tokens=400,
    response_format={"type": "json_object"},
    messages=[
        {"role": "system", "content": 'Reply with a JSON object with the keys "language" and "summary".'},
        {"role": "user", "content": "def add(a, b): return a + b"},
    ],
)

choice = resp.choices[0]
if choice.finish_reason == "length":
    raise RuntimeError("Cut off by max_tokens: the JSON is incomplete")
data = json.loads(choice.message.content)
print(data)
```

```typescript title="Node.js"
import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://tokens.bd/v1", apiKey: process.env.TOKENS_API_KEY });

const resp = await client.chat.completions.create({
  model: "deepseek/deepseek-v4.1-flash",
  max_tokens: 400,
  response_format: { type: "json_object" },
  messages: [
    { role: "system", content: 'Reply with a JSON object with the keys "language" and "summary".' },
    { role: "user", content: "def add(a, b): return a + b" },
  ],
});

const choice = resp.choices[0];
if (choice.finish_reason === "length") throw new Error("Cut off by max_tokens: the JSON is incomplete");
console.log(JSON.parse(choice.message.content ?? "{}"));
```

:::

parse করার আগে `finish_reason` দেখে নিন। সেটা `length` হলে বুঝবেন model object-এর মাঝখানে `max_tokens`-এ ঠেকে গেছে, তাই JSON কাটা পড়েছে। reasoning model ওই limit-এর একটা অংশ আগে ভাবতেই খরচ করে, তাই তাদের বেশি জায়গা দিন। [Reasoning ও thinking model](/docs/reasoning) পেজ দেখুন।

## response_format-এ JSON Schema

JSON Schema output-এ উত্তর আপনার দেওয়া schema-র ভেতরে বাঁধা থাকে। Chat Completions-এ request দেখতে এরকম (shape-টা OpenAI-র, October 2026-এ দেখা হয়েছে [OpenAI-এর structured outputs guide](https://developers.openai.com/api/docs/guides/structured-outputs) আর [migration guide](https://developers.openai.com/api/docs/guides/migrate-to-responses) থেকে):

```json
{
  "response_format": {
    "type": "json_schema",
    "json_schema": {
      "name": "ticket",
      "strict": true,
      "schema": {
        "type": "object",
        "properties": {
          "title": { "type": "string" },
          "severity": { "type": "string", "enum": ["low", "medium", "high"] },
          "needs_followup": { "type": "boolean" }
        },
        "required": ["title", "severity", "needs_followup"],
        "additionalProperties": false
      }
    }
  }
}
```

`strict: true` দিলে OpenAI schema-র ওপর এই নিয়মগুলো খাটায়, আর নিয়ম ভাঙলে schema ফিরিয়ে দেয়:

- Root অবশ্যই object হতে হবে, আর সেটা `anyOf` হতে পারবে না।
- প্রতিটা property `required`-এ থাকতে হবে। কোনো field optional করতে চাইলে `null` মঞ্জুর করুন: `{"type": ["string", "null"]}`।
- প্রতিটা object-এ `"additionalProperties": false` লাগবে।
- সব মিলিয়ে সর্বোচ্চ 5,000টা property আর 10 স্তর nesting।

অন্য maker-রা JSON Schema-র একটা অংশ support করে, আর সেই অংশটা একেক জায়গায় একেক রকম। provider আপনার schema ফিরিয়ে দিলে সেটা সহজ করুন: সমতল object, সাধারণ type, আর নির্দিষ্ট choice-এর জন্য `enum`।

:::code-tabs

```python title="Python"
import json
import os
from openai import OpenAI

client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"])

schema = {
    "type": "object",
    "properties": {
        "title": {"type": "string"},
        "severity": {"type": "string", "enum": ["low", "medium", "high"]},
        "needs_followup": {"type": "boolean"},
    },
    "required": ["title", "severity", "needs_followup"],
    "additionalProperties": False,
}

resp = client.chat.completions.create(
    model="deepseek/deepseek-v4.1-flash",
    max_tokens=400,
    messages=[
        {"role": "user", "content": "Users report that the login page returns a 500 after the last deploy."}
    ],
    response_format={
        "type": "json_schema",
        "json_schema": {"name": "ticket", "strict": True, "schema": schema},
    },
)

choice = resp.choices[0]
if getattr(choice.message, "refusal", None):
    print("The model refused:", choice.message.refusal)
elif choice.finish_reason != "stop":
    print("Incomplete answer, finish_reason =", choice.finish_reason)
else:
    ticket = json.loads(choice.message.content)
    print(ticket["severity"], ticket["title"])
```

```typescript title="Node.js"
import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://tokens.bd/v1", apiKey: process.env.TOKENS_API_KEY });

const schema = {
  type: "object",
  properties: {
    title: { type: "string" },
    severity: { type: "string", enum: ["low", "medium", "high"] },
    needs_followup: { type: "boolean" },
  },
  required: ["title", "severity", "needs_followup"],
  additionalProperties: false,
};

const resp = await client.chat.completions.create({
  model: "deepseek/deepseek-v4.1-flash",
  max_tokens: 400,
  messages: [
    { role: "user", content: "Users report that the login page returns a 500 after the last deploy." },
  ],
  response_format: { type: "json_schema", json_schema: { name: "ticket", strict: true, schema } },
});

const choice = resp.choices[0];
if (choice.finish_reason !== "stop") {
  console.log("Incomplete or refused answer:", choice.finish_reason, choice.message.refusal);
} else {
  console.log(JSON.parse(choice.message.content ?? "{}"));
}
```

```bash title="cURL"
curl https://tokens.bd/v1/chat/completions \
  -H "Authorization: Bearer $TOKENS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek/deepseek-v4.1-flash",
    "max_tokens": 400,
    "messages": [
      {"role": "user", "content": "Users report that the login page returns a 500 after the last deploy."}
    ],
    "response_format": {
      "type": "json_schema",
      "json_schema": {
        "name": "ticket",
        "strict": true,
        "schema": {
          "type": "object",
          "properties": {
            "title": {"type": "string"},
            "severity": {"type": "string", "enum": ["low", "medium", "high"]},
            "needs_followup": {"type": "boolean"}
          },
          "required": ["title", "severity", "needs_followup"],
          "additionalProperties": false
        }
      }
    }
  }'
```

:::

parse করার আগে তিনটা ফলাফল সামলে নিন: refusal (OpenAI এটা `message.refusal`-এ জানায), অসম্পূর্ণ উত্তর (`finish_reason` হলো `length` বা `content_filter`), আর এমন উত্তর যা parse হয় কিন্তু আপনার নিজের check পাস করে না। OpenAI আর Anthropic-এর SDK-তে এমন helper আছে যা Pydantic model বা Zod schema থেকে schema বানিয়ে দেয় আর উত্তর parse করে। তারা উপরের মতো একই JSON পাঠায়, তাই model feature-টা support করলে Tokens দিয়েও চলে।

:::note[Responses API]
Responses API একই schema নেয় `response_format`-এর বদলে `text.format`-এর নিচে, যেখানে `name`, `strict` আর `schema` এক ধাপ ওপরে থাকে, ভেতরে আলাদা `json_schema` key থাকে না। [Responses API](/docs/responses) দেখুন। Support নির্ভর করে model-এর পেছনের provider-এর ওপর।
:::

## Tools দিয়ে structured output

forced tool call সবচেয়ে বেশি জায়গায় চলে, কারণ `response_format`-এর চেয়ে অনেক বেশি model tool calling support করে, আর এটা gateway-র format translation-এও টিকে যায়। একটা function বানান যার `parameters` আপনার চাওয়া schema, model-কে সেটা call করতে বাধ্য করুন, তারপর arguments পড়ে নিন।

```python title="Chat Completions: forced function"
import json
import os
from openai import OpenAI

client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"])

tools = [
    {
        "type": "function",
        "function": {
            "name": "record_ticket",
            "description": "Record a support ticket extracted from the user's message.",
            "strict": True,
            "parameters": {
                "type": "object",
                "properties": {
                    "title": {"type": "string"},
                    "severity": {"type": "string", "enum": ["low", "medium", "high"]},
                },
                "required": ["title", "severity"],
                "additionalProperties": False,
            },
        },
    }
]

resp = client.chat.completions.create(
    model="deepseek/deepseek-v4.1-flash",
    max_tokens=400,
    messages=[{"role": "user", "content": "The login page returns a 500 after the last deploy."}],
    tools=tools,
    tool_choice={"type": "function", "function": {"name": "record_ticket"}},
)

call = resp.choices[0].message.tool_calls[0]
ticket = json.loads(call.function.arguments)
print(ticket)
```

Chat Completions-এ `arguments` একটা JSON string, তাই এটা parse করতে হবে। `function` object-এর ভেতরের `strict` হলো schema হুবহু মানানোর জন্য OpenAI-র switch, আর এর schema-র নিয়ম উপরের মতোই। যেসব model ও provider এটা চেনে না, তারা ignore করতে পারে। এখানে কিছুই Tokens চালায় না, Tokens শুধু call-টা বয়ে নিয়ে যায়। নির্দিষ্ট একটা function-এ `tool_choice` বেঁধে দেওয়ার support model ভেদে আলাদা: কেউ ignore করে, কেউ `400` দেয়। [Tool calling](/docs/tool-calling) দেখুন।

OpenAI-র documentation বলছে, GPT-5.4 আর তার পরের model-এ `reasoning_effort` `none` ছাড়া অন্য কিছু হলে Chat Completions tool calling support করে না। কোনো OpenAI reasoning model-এ forced tool call fail করলে `reasoning_effort` `none` করে দিন, অথবা Responses API ব্যবহার করুন।

Messages-এ একই ধারণা চলে Anthropic-এর tool format-এ, আর arguments `tool_use` block-এ আগে থেকেই object হয়ে parse করা অবস্থায় ফেরত আসে:

```python title="Messages: forced tool"
import os
import anthropic

client = anthropic.Anthropic(base_url="https://tokens.bd", api_key=os.environ["TOKENS_API_KEY"])

tools = [
    {
        "name": "record_ticket",
        "description": "Record a support ticket extracted from the user's message.",
        "input_schema": {
            "type": "object",
            "properties": {
                "title": {"type": "string"},
                "severity": {"type": "string", "enum": ["low", "medium", "high"]},
            },
            "required": ["title", "severity"],
        },
    }
]

message = client.messages.create(
    model="deepseek/deepseek-v4.1-flash",
    max_tokens=400,
    tools=tools,
    tool_choice={"type": "tool", "name": "record_ticket"},
    messages=[{"role": "user", "content": "The login page returns a 500 after the last deploy."}],
)

block = next(b for b in message.content if b.type == "tool_use")
print(block.input)
```

Anthropic জানায়, `tool_choice`-এ `any` বা `tool` দিয়ে tool জোর করলে তাকে তাদের manual extended thinking (`thinking.type: "enabled"`)-এর সাথে একসাথে চালানো যায় না, আর তাদের একেবারে নতুন কয়েকটা model adaptive thinking থাকলেও forced tool use নেয় না। Claude model-এ thinking আর structured output দুটোই চাইলে আগে [Reasoning ও thinking model](/docs/reasoning) পড়ুন।

## Messages-এ output_config দিয়ে JSON Schema

Anthropic-এর Messages API-র নিজস্ব JSON Schema output আছে, যা `output_config.format`-এ দিতে হয় (October 2026-এ দেখা হয়েছে [Anthropic-এর structured outputs guide](https://platform.claude.com/docs/en/build-with-claude/structured-outputs) থেকে)। উত্তর আসে `text` block-এর ভেতর text হিসেবে, সেই text-টাই JSON। Anthropic tool definition-এ `strict: true`-ও দেয়, যাতে tool-এর নাম আর input schema-র সাথে মিলে যাওয়া নিশ্চিত হয়।

```bash title="cURL"
curl https://tokens.bd/v1/messages \
  -H "x-api-key: $TOKENS_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{
    "model": "deepseek/deepseek-v4.1-flash",
    "max_tokens": 400,
    "messages": [
      {"role": "user", "content": "The login page returns a 500 after the last deploy."}
    ],
    "output_config": {
      "format": {
        "type": "json_schema",
        "schema": {
          "type": "object",
          "properties": {
            "title": {"type": "string"},
            "severity": {"type": "string", "enum": ["low", "medium", "high"]}
          },
          "required": ["title", "severity"],
          "additionalProperties": false
        }
      }
    }
  }'
```

```python title="Python"
import json
import os
import anthropic

client = anthropic.Anthropic(base_url="https://tokens.bd", api_key=os.environ["TOKENS_API_KEY"])

response = client.messages.create(
    model="deepseek/deepseek-v4.1-flash",
    max_tokens=400,
    messages=[{"role": "user", "content": "The login page returns a 500 after the last deploy."}],
    output_config={
        "format": {
            "type": "json_schema",
            "schema": {
                "type": "object",
                "properties": {
                    "title": {"type": "string"},
                    "severity": {"type": "string", "enum": ["low", "medium", "high"]},
                },
                "required": ["title", "severity"],
                "additionalProperties": False,
            },
        }
    },
)

if response.stop_reason in ("refusal", "max_tokens"):
    raise RuntimeError(f"No valid JSON: stop_reason = {response.stop_reason}")
text = next(b.text for b in response.content if b.type == "text")
print(json.loads(text))
```

এর জন্য Anthropic SDK-র নতুন version লাগে, যে version `output_config` চেনে। আপনার SDK argument-টা ফিরিয়ে দিলে SDK upgrade করুন। Anthropic কোন কোন Claude model-এ feature-টা চলে তার তালিকা দেয়, আর JSON Schema-র শুধু একটা অংশ support করে: recursive schema চলে না, সংখ্যা বা string-এর দৈর্ঘ্যের constraint চলে না, object-এ `additionalProperties: false` দিতে হয়, আর unsupported feature দিলে `400` আসে। `stop_reason` `refusal` হলেও `200` আসে আর bill হয়, তখন output schema না-ও মানতে পারে; `max_tokens`-এ ঠেকলে উত্তর অসম্পূর্ণ থাকতে পারে।

Tokens-এ এই পথের সীমাটা মনে রাখুন। model-টা যে provider-এর কাছে পাওয়া যায় সে Anthropic format বুঝলে এটা কাজ করে। gateway আপনার Messages request OpenAI format-এ translate করলে `output_config` বাদ পড়ে যায় (উপরে দেখুন)। সব model-এর জন্য একটাই code path চাইলে forced-tool পদ্ধতি ব্যবহার করুন।

## Streaming আর structured output

`"stream": true` দিলেও structured output চলে, কিন্তু JSON valid হয় stream পুরো শেষ হলে। `content` delta-গুলো (বা tool-call `arguments`-এর টুকরো, বা Anthropic-এর `input_json_delta` অংশগুলো) জোড়া দিন, আর শেষে একবার parse করুন। অর্ধেক text parse করতে যাবেন না। [streaming](/docs/streaming) দেখুন।

## খরচ আর limit

- schema, tool definition আর আপনার instruction প্রতিটা request-এ input token হিসেবে গোনা হয়। বড় schema প্রতিটা call-এ খরচ বাড়ায়।
- Tokens provider-এর report করা usage-ই bill করে, model-এর input আর output price ধরে। structured output-এর জন্য বাড়তি কোনো চার্জ নেই।
- `max_tokens` পুরো উত্তরের সীমা। কাটা পড়া JSON object parse হয় না, তাই আপনার সবচেয়ে লম্বা সম্ভাব্য উত্তরের চেয়ে বেশি রাখুন, আর model আগে ভাবলে আরও জায়গা যোগ করুন।
- request forward করার আগে gateway `max_tokens` ধরে worst-case খরচ reserve করে। reservation নিয়ে [Chat Completions](/docs/chat-completions) পেজ দেখুন।

## নির্ভরযোগ্য করার উপায়

- strict mode ব্যবহার করলেও প্রতিটা উত্তর নিজের code-এ নিজের schema-র সাথে validate করুন (Pydantic, Zod, `jsonschema`)।
- schema সমতল রাখুন আর প্রতিটা field-এ ছোট একটা `description` দিন; model সেটা পড়ে।
- extraction-এর কাজে `temperature` কম রাখুন। কিছু reasoning model এটা মানে না বা ignore করে; [Reasoning ও thinking model](/docs/reasoning) দেখুন।
- parse fail করলে error message জুড়ে একবার retry করুন, তারপর স্পষ্টভাবে fail করুন। প্রতিটা retry নতুন bill-হওয়া request।
- production-এ যে model চালাবেন, ঠিক সেটাতেই test করুন। এক model-এ চলা schema আরেক model ফিরিয়ে দিতে পারে।

---
Page: https://tokens.bd/bn/docs/structured-output
