POST https://tokens.bd/v1/chat/completions হলো OpenAI-compatible chat completions API। বেশির ভাগ SDK আর coding agent এই endpoint-ই ব্যবহার করে। Request আর response OpenAI-র format-এই চলে। Gateway আগে আপনার key, plan আর limit যাচাই করে, তারপর body-টা আপনার দেওয়া model-এর upstream provider-এর কাছে পাঠিয়ে দেয়।
Chat completions request পাঠান#
curl https://tokens.bd/v1/chat/completions \
-H "Authorization: Bearer $TOKENS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek/deepseek-v4.1-flash",
"messages": [
{"role": "system", "content": "You are a concise senior engineer."},
{"role": "user", "content": "What does HTTP 429 mean?"}
],
"temperature": 0.2,
"max_tokens": 300
}'import os
from openai import OpenAI
client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"])
resp = client.chat.completions.create(
model="deepseek/deepseek-v4.1-flash",
messages=[
{"role": "system", "content": "You are a concise senior engineer."},
{"role": "user", "content": "What does HTTP 429 mean?"},
],
temperature=0.2,
max_tokens=300,
)
print(resp.choices[0].message.content)
print(resp.usage)import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://tokens.bd/v1", apiKey: process.env.TOKENS_API_KEY });
const resp = await client.chat.completions.create({
model: "deepseek/deepseek-v4.1-flash",
messages: [
{ role: "system", content: "You are a concise senior engineer." },
{ role: "user", content: "What does HTTP 429 mean?" },
],
temperature: 0.2,
max_tokens: 300,
});
console.log(resp.choices[0].message.content, resp.usage);Content-Type: application/json header-টা বাদ দেবেন না। এটা না থাকলে gateway body পড়তে পারে না, আর model field দেওয়া থাকলেও 400 error আসে, যেখানে বলা হয় request-এ model field থাকতে হবে।
Request-এর field#
| Field | Type | নোট |
|---|---|---|
model | string | আবশ্যক। Catalog id, যেমন deepseek/deepseek-v4.1-flash। আপনার key-র model-এর list পাবেন GET /v1/models-এ। |
messages | array | আবশ্যক। প্রতিটা object-এ role (system, user, assistant, tool) আর content থাকবে। |
temperature | number | সাধারণত 0 থেকে 2। কিছু reasoning model এটা উপেক্ষা করে বা reject করে। |
max_tokens | integer | কতগুলো token পর্যন্ত generate হবে তার সীমা। এটা সবসময় দিন; কেন, তা নিচে reservation-এর নোটে আছে। |
max_completion_tokens | integer | একই সীমার জন্য OpenAI-র নতুন নাম। কিছু model-এ max_tokens-এর বদলে এটাই দিতে হয়। |
stream | boolean | true দিলে Server-Sent Events আসে। দেখুন streaming। |
stream_options.include_usage | boolean | stream: true-র সঙ্গে দিলে শেষে usage-সহ একটা বাড়তি chunk আসে। না চাইলে আসে না। |
tools | array | Function-এর definition। দেখুন tool calling। |
tool_choice | string বা object | "auto", "none", "required", বা নির্দিষ্ট একটা function। Model ভেদে support আলাদা। |
response_format | object | {"type": "json_object"} বা json_schema object, যেসব model support করে তাদের জন্য। |
n | integer | কয়টা choice চান, 1 থেকে 4। এর বাইরে দিলে 400 invalid_request আসে। |
stop, top_p, seed, presence_penalty, frequency_penalty | নানা ধরনের | যেমন পাঠান তেমনই পাস হয়ে যায়। |
Parameter নির্ভর করে upstream model-এর ওপর
model আর n ছাড়া বাকি field-গুলো gateway যাচাই করে না, সরাসরি model-এর provider-এর কাছে পাঠিয়ে দেয়। তবে দুটো ব্যতিক্রম আছে। প্রথমত, OpenAI-র reasoning model-এর (o-series আর GPT-5 ও তার পরের model) ক্ষেত্রে gateway max_tokens-কে max_completion_tokens নামে পাঠায়, আর temperature ও top_p 1 না হলে সরিয়ে দেয়, কারণ ওই model-গুলো এই দুটো field reject করে। দ্বিতীয়ত, model যদি শুধু Anthropic-এর provider থেকে আসে, তাহলে ওই format-এ যে field-এর জায়গা নেই সেগুলো বাদ পড়ে: response_format, reasoning_effort, n, seed, logprobs আর penalty field-গুলো। বিস্তারিত দেখুন Structured output আর Reasoning পেজে। কোনো model যদি response_format, tools বা আপনার দেওয়া temperature support না করে, তাহলে কী হবে সেটা provider ঠিক করে: field উপেক্ষাও করতে পারে, request reject-ও করতে পারে। কোনো feature-এর ওপর ভরসা করার আগে catalog-এ model-এর পেজটা দেখে নিন।
Request body সর্বোচ্চ 10 MB হতে পারে। এর চেয়ে বড় হলে 413 আসে।
Response-এর উদাহরণ#
{
"id": "chatcmpl-a1b2c3",
"object": "chat.completion",
"created": 1790000000,
"model": "deepseek/deepseek-v4.1-flash",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "429 Too Many Requests: the server is rate limiting you. Back off and retry after the Retry-After interval."
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 27,
"completion_tokens": 24,
"total_tokens": 51
}
}Body-টা আসে upstream provider থেকে। তাই id-র format, model-এর string আর বাড়তি field (যেমন reasoning_content বা system_fingerprint) model ভেদে আলাদা হতে পারে।
usage object#
usage-এ থাকে provider যা গুনেছে:
| Field | মানে |
|---|---|
prompt_tokens | Input token, cache থেকে আসাগুলোসহ |
completion_tokens | Output token; যেসব model reasoning token এভাবে হিসাব করে, তাদের ক্ষেত্রে সেগুলোসহ |
total_tokens | দুটোর যোগফল |
prompt_tokens_details.cached_tokens | Provider-এর cache থেকে আসা input token, যদি provider জানায় |
Billing হয় upstream-এর এই হিসাব ধরে। Cache-read rate আছে এমন model-এ cache থেকে আসা input সেই rate-এ ধরা হয়। কোনো provider যদি usage একদমই না পাঠায়, gateway request আর response-এর size দেখে আন্দাজ করে নেয়। প্রতিটা request-এর খরচ দেখবেন usage analytics-এ; model ধরে rate আছে model catalog আর pricing পেজে।
max_tokens কীভাবে admission-এ প্রভাব ফেলে#
Request পাঠানোর আগে gateway আপনার plan-এর credit বা Wallet থেকে সবচেয়ে খারাপ ক্ষেত্রের খরচটা reserve করে রাখে। Output-এর অংশ ধরা হয় max_tokens (বা max_completion_tokens) দিয়ে। এটা না দিলে reservation-এ ধরে নেওয়া হয় 8,192 output token। এর দুটো বাস্তব ফল:
- Key-র monthly spend cap-এর কাছাকাছি থাকলে বড়
max_tokens-এ request ফিরিয়ে দেওয়া হতে পারে, অথচ ছোটটা চলে যায়। - আপনার ব্যালান্স দিয়ে চাওয়া
max_tokensপোষালে না হলে gateway সেটা কমিয়ে ব্যালান্সে যতটা কুলায় ততটায় নামিয়ে আনতে পারে (16-র নিচে নামে না)। তখন ছোট একটা উত্তরের সঙ্গেfinish_reason: "length"দেখবেন। Billing-এ গিয়ে টাকা যোগ করুন, নয়তো নিজেইmax_tokensকমিয়ে দিন।
চার্জ হয় আসল usage অনুযায়ী, reservation অনুযায়ী নয়।
Error#
Gateway-র error OpenAI-র error shape-এ আসে, সঙ্গে একটা code থাকে যা দেখে আপনি code-এ branch করতে পারেন। প্রতিটা response-এ x-tokens-request-id header থাকে। পুরো table আছে errors পেজে, আর প্রতি মিনিটের ও concurrency limit আছে rate limits পেজে।