Structured output means asking a model for JSON your code can parse, instead of prose. There are four ways to do it through Tokens, and which one works depends on the model and on how the gateway reaches it:
| Method | Endpoint | How strong the guarantee is |
|---|---|---|
JSON mode: response_format: {"type": "json_object"} | Chat Completions | Valid JSON, but no promise about the keys |
JSON Schema: response_format: {"type": "json_schema", ...} | Chat Completions | Output follows your schema (with strict: true), if the model supports it |
| A forced tool call whose arguments are your schema | Chat Completions, Messages | Arguments follow the schema as well as the model manages; works on more models |
output_config.format with a JSON Schema | Messages | Output follows your schema, on models that support it |
None of this is enforced by Tokens. The gateway does not read, validate or repair the JSON. It forwards the request fields and returns the provider's answer, so the guarantees above are the provider's, and a model that does not support a feature either ignores the field or rejects the request. Always parse and validate the result in your own code.
What the gateway does and drops#
- On Chat Completions to a provider that speaks the OpenAI format,
response_formatis forwarded unchanged. - On Messages to a provider that speaks the Anthropic format,
output_config,toolsandtool_choiceare forwarded unchanged. - When a model is only available from a provider that speaks the other format, the gateway translates the request, and the translation carries only a fixed list of fields. A Chat Completions request to an Anthropic-format provider does not carry
response_format(norn,seed,logprobsor the penalty fields). A Messages request to an OpenAI-format provider does not carryoutput_config. In both cases the field has no effect and nothing tells you so; the model just answers in free text. - Tools survive translation in both directions, which is why a forced tool call is the most portable method. See Tool calling and Messages.
If a structured request comes back as prose, the likely causes are a model without support, or a translated path that dropped the field. Try the tool-call method on the same model.
How to find out whether a model supports it#
The model catalog at /models records context window, prices and a description, but no field for JSON mode or JSON Schema support. To check:
- Read the maker's documentation for the model (for example, whether it lists "JSON output" or "structured outputs").
- Send a small request with the field and look at the result: a 400 means the model or provider rejects it, free-form prose means the field was ignored, valid JSON means it works for that request.
- Run the test several times. Unconstrained JSON mode can pass once and fail on the next prompt.
Choosing a model lists the models Tokens compares most often.
JSON mode#
JSON mode makes the model return a valid JSON object, with no guarantee about which keys it contains. Tell the model in a message to produce JSON and describe the shape you want. OpenAI's documentation says you must instruct the model to produce JSON in some message in the conversation; without that the model can fill the response with whitespace until it hits the token limit, and the API can reject the request when the word "JSON" is missing from the context.
curl https://tokens.bd/v1/chat/completions \
-H "Authorization: Bearer $TOKENS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek/deepseek-v4.1-flash",
"max_tokens": 400,
"response_format": {"type": "json_object"},
"messages": [
{"role": "system", "content": "Reply with a JSON object with the keys \"language\" and \"summary\"."},
{"role": "user", "content": "def add(a, b): return a + b"}
]
}'import json
import os
from openai import OpenAI
client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"])
resp = client.chat.completions.create(
model="deepseek/deepseek-v4.1-flash",
max_tokens=400,
response_format={"type": "json_object"},
messages=[
{"role": "system", "content": 'Reply with a JSON object with the keys "language" and "summary".'},
{"role": "user", "content": "def add(a, b): return a + b"},
],
)
choice = resp.choices[0]
if choice.finish_reason == "length":
raise RuntimeError("Cut off by max_tokens: the JSON is incomplete")
data = json.loads(choice.message.content)
print(data)import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://tokens.bd/v1", apiKey: process.env.TOKENS_API_KEY });
const resp = await client.chat.completions.create({
model: "deepseek/deepseek-v4.1-flash",
max_tokens: 400,
response_format: { type: "json_object" },
messages: [
{ role: "system", content: 'Reply with a JSON object with the keys "language" and "summary".' },
{ role: "user", content: "def add(a, b): return a + b" },
],
});
const choice = resp.choices[0];
if (choice.finish_reason === "length") throw new Error("Cut off by max_tokens: the JSON is incomplete");
console.log(JSON.parse(choice.message.content ?? "{}"));Check finish_reason before parsing. If it is length, the model hit max_tokens in the middle of the object and the JSON is cut off. Reasoning models spend part of that limit on thinking first, so give them more room; see Reasoning and thinking models.
JSON Schema with response_format#
JSON Schema output constrains the answer to a schema you supply. On Chat Completions the request looks like this (the shape is OpenAI's, checked October 2026 in OpenAI's structured outputs guide and migration guide):
{
"response_format": {
"type": "json_schema",
"json_schema": {
"name": "ticket",
"strict": true,
"schema": {
"type": "object",
"properties": {
"title": { "type": "string" },
"severity": { "type": "string", "enum": ["low", "medium", "high"] },
"needs_followup": { "type": "boolean" }
},
"required": ["title", "severity", "needs_followup"],
"additionalProperties": false
}
}
}
}With strict: true OpenAI applies these rules to the schema, and rejects a schema that breaks them:
- The root must be an object, and it cannot be an
anyOf. - Every property must be listed in
required. To make a field optional, allownull:{"type": ["string", "null"]}. - Every object needs
"additionalProperties": false. - At most 5,000 properties in total and 10 levels of nesting.
Other makers support a subset of JSON Schema, and the subset differs. If a provider rejects your schema, simplify it: flat objects, simple types, enum for fixed choices.
import json
import os
from openai import OpenAI
client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"])
schema = {
"type": "object",
"properties": {
"title": {"type": "string"},
"severity": {"type": "string", "enum": ["low", "medium", "high"]},
"needs_followup": {"type": "boolean"},
},
"required": ["title", "severity", "needs_followup"],
"additionalProperties": False,
}
resp = client.chat.completions.create(
model="deepseek/deepseek-v4.1-flash",
max_tokens=400,
messages=[
{"role": "user", "content": "Users report that the login page returns a 500 after the last deploy."}
],
response_format={
"type": "json_schema",
"json_schema": {"name": "ticket", "strict": True, "schema": schema},
},
)
choice = resp.choices[0]
if getattr(choice.message, "refusal", None):
print("The model refused:", choice.message.refusal)
elif choice.finish_reason != "stop":
print("Incomplete answer, finish_reason =", choice.finish_reason)
else:
ticket = json.loads(choice.message.content)
print(ticket["severity"], ticket["title"])import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://tokens.bd/v1", apiKey: process.env.TOKENS_API_KEY });
const schema = {
type: "object",
properties: {
title: { type: "string" },
severity: { type: "string", enum: ["low", "medium", "high"] },
needs_followup: { type: "boolean" },
},
required: ["title", "severity", "needs_followup"],
additionalProperties: false,
};
const resp = await client.chat.completions.create({
model: "deepseek/deepseek-v4.1-flash",
max_tokens: 400,
messages: [
{ role: "user", content: "Users report that the login page returns a 500 after the last deploy." },
],
response_format: { type: "json_schema", json_schema: { name: "ticket", strict: true, schema } },
});
const choice = resp.choices[0];
if (choice.finish_reason !== "stop") {
console.log("Incomplete or refused answer:", choice.finish_reason, choice.message.refusal);
} else {
console.log(JSON.parse(choice.message.content ?? "{}"));
}curl https://tokens.bd/v1/chat/completions \
-H "Authorization: Bearer $TOKENS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek/deepseek-v4.1-flash",
"max_tokens": 400,
"messages": [
{"role": "user", "content": "Users report that the login page returns a 500 after the last deploy."}
],
"response_format": {
"type": "json_schema",
"json_schema": {
"name": "ticket",
"strict": true,
"schema": {
"type": "object",
"properties": {
"title": {"type": "string"},
"severity": {"type": "string", "enum": ["low", "medium", "high"]},
"needs_followup": {"type": "boolean"}
},
"required": ["title", "severity", "needs_followup"],
"additionalProperties": false
}
}
}
}'Handle three outcomes before you parse: a refusal (OpenAI reports it in message.refusal), an incomplete answer (finish_reason of length or content_filter), and an answer that parses but fails your own checks. The OpenAI and Anthropic SDKs have helpers that build the schema from a Pydantic model or a Zod schema and parse the reply; they send the same JSON as above, so they work through Tokens when the model supports the feature.
Responses API
The Responses API takes the same schema under text.format instead of response_format, with name, strict and schema one level up and no nested json_schema key. See Responses API. Support depends on the provider behind the model.
Structured output with tools#
A forced tool call is the most portable method, because tool calling is supported by many more models than response_format, and it survives the gateway's format translation. Define one function whose parameters are the schema you want, force the model to call it, and read the arguments.
import json
import os
from openai import OpenAI
client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"])
tools = [
{
"type": "function",
"function": {
"name": "record_ticket",
"description": "Record a support ticket extracted from the user's message.",
"strict": True,
"parameters": {
"type": "object",
"properties": {
"title": {"type": "string"},
"severity": {"type": "string", "enum": ["low", "medium", "high"]},
},
"required": ["title", "severity"],
"additionalProperties": False,
},
},
}
]
resp = client.chat.completions.create(
model="deepseek/deepseek-v4.1-flash",
max_tokens=400,
messages=[{"role": "user", "content": "The login page returns a 500 after the last deploy."}],
tools=tools,
tool_choice={"type": "function", "function": {"name": "record_ticket"}},
)
call = resp.choices[0].message.tool_calls[0]
ticket = json.loads(call.function.arguments)
print(ticket)On Chat Completions arguments is a JSON string, so parse it. strict inside the function object is OpenAI's switch for exact schema adherence and has the same schema rules as above; models and providers that do not know it may ignore it. Nothing here is run by Tokens, which only carries the call. Support for tool_choice set to a specific function varies by model: some ignore it, some return a 400. See Tool calling.
OpenAI's documentation says Chat Completions does not support tool calling with a reasoning_effort other than none on GPT-5.4 and later models. If a forced tool call fails on an OpenAI reasoning model, set reasoning_effort to none or use the Responses API.
On Messages the same idea uses Anthropic's tool format, and the arguments come back already parsed as an object in a tool_use block:
import os
import anthropic
client = anthropic.Anthropic(base_url="https://tokens.bd", api_key=os.environ["TOKENS_API_KEY"])
tools = [
{
"name": "record_ticket",
"description": "Record a support ticket extracted from the user's message.",
"input_schema": {
"type": "object",
"properties": {
"title": {"type": "string"},
"severity": {"type": "string", "enum": ["low", "medium", "high"]},
},
"required": ["title", "severity"],
},
}
]
message = client.messages.create(
model="deepseek/deepseek-v4.1-flash",
max_tokens=400,
tools=tools,
tool_choice={"type": "tool", "name": "record_ticket"},
messages=[{"role": "user", "content": "The login page returns a 500 after the last deploy."}],
)
block = next(b for b in message.content if b.type == "tool_use")
print(block.input)Anthropic notes that forcing a tool with tool_choice of any or tool cannot be combined with its manual extended thinking (thinking.type: "enabled"), and that a few of its newest models do not accept forced tool use even with adaptive thinking. If you need both thinking and structured output on Claude models, read Reasoning and thinking models first.
JSON Schema on Messages with output_config#
Anthropic's Messages API has its own JSON Schema output, set in output_config.format (checked October 2026 in Anthropic's structured outputs guide). The reply is text in a text block that holds the JSON. Anthropic also offers strict: true on a tool definition to guarantee that tool names and inputs match the schema.
curl https://tokens.bd/v1/messages \
-H "x-api-key: $TOKENS_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "deepseek/deepseek-v4.1-flash",
"max_tokens": 400,
"messages": [
{"role": "user", "content": "The login page returns a 500 after the last deploy."}
],
"output_config": {
"format": {
"type": "json_schema",
"schema": {
"type": "object",
"properties": {
"title": {"type": "string"},
"severity": {"type": "string", "enum": ["low", "medium", "high"]}
},
"required": ["title", "severity"],
"additionalProperties": false
}
}
}
}'import json
import os
import anthropic
client = anthropic.Anthropic(base_url="https://tokens.bd", api_key=os.environ["TOKENS_API_KEY"])
response = client.messages.create(
model="deepseek/deepseek-v4.1-flash",
max_tokens=400,
messages=[{"role": "user", "content": "The login page returns a 500 after the last deploy."}],
output_config={
"format": {
"type": "json_schema",
"schema": {
"type": "object",
"properties": {
"title": {"type": "string"},
"severity": {"type": "string", "enum": ["low", "medium", "high"]},
},
"required": ["title", "severity"],
"additionalProperties": False,
},
}
},
)
if response.stop_reason in ("refusal", "max_tokens"):
raise RuntimeError(f"No valid JSON: stop_reason = {response.stop_reason}")
text = next(b.text for b in response.content if b.type == "text")
print(json.loads(text))This needs a recent Anthropic SDK that knows output_config; if your SDK version rejects the argument, upgrade it. Anthropic lists the Claude models that support the feature, and it supports only a subset of JSON Schema: no recursive schemas, no numeric or string length constraints, additionalProperties: false on objects, and unsupported features return a 400. A stop_reason of refusal still returns a 200 and is billed, and the output may not match the schema; max_tokens can leave it incomplete.
Remember the limits of this path on Tokens. It works when the model is served by a provider that speaks the Anthropic format. When the gateway has to translate your Messages request to the OpenAI format, output_config is dropped (see above). Use the forced-tool method if you need one code path for every model.
Streaming and structured output#
Structured output works with "stream": true, but the JSON is only valid when the stream is complete. Concatenate the content deltas (or the tool-call arguments fragments, or Anthropic's input_json_delta pieces) and parse once at the end. Do not parse partial text. See streaming.
Cost and limits#
- The schema, tool definitions and your instructions are input tokens on every request. A large schema costs on every call.
- Tokens bills the usage the provider reports, at the model's input and output prices. There is no surcharge for structured output.
max_tokensbounds the whole answer. A truncated JSON object does not parse, so set it above the longest answer you expect, and add room if the model reasons first.- The gateway reserves the worst-case cost from
max_tokensbefore it forwards the request. See the reservation notes in Chat Completions.
Make it reliable#
- Validate every reply against your own schema in code (Pydantic, Zod,
jsonschema), even when you used strict mode. - Keep schemas flat and give each field a short
description; the model reads it. - Set low
temperaturefor extraction tasks. Some reasoning models reject or ignore it; see Reasoning and thinking models. - On a parse failure, retry once with the error message appended, then fail loudly. Each retry is a new billed request.
- Test the exact model you will run in production. A schema that works on one model can be rejected by another.