POST https://tokens.bd/v1/completions is the original OpenAI text completions endpoint: you send a prompt string and the model continues it. Tokens passes it through for older tools and scripts that still call it. For anything new, use chat completions, which every chat model supports and which most coding agents expect.
Not every model serves this endpoint
Tokens forwards the request to the model's provider, and the provider decides whether the model can do plain text completion. Many chat models can't, and there is no list of the ones that can. Test your model with a small request (below) before you build on it.
Make a completions request#
curl https://tokens.bd/v1/completions \
-H "Authorization: Bearer $TOKENS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek/deepseek-v4.1-flash",
"prompt": "A one-line definition of HTTP 429:",
"max_tokens": 40,
"temperature": 0
}'import os
from openai import OpenAI
client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"])
resp = client.completions.create(
model="deepseek/deepseek-v4.1-flash",
prompt="A one-line definition of HTTP 429:",
max_tokens=40,
temperature=0,
)
print(resp.choices[0].text)
print(resp.usage)import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://tokens.bd/v1", apiKey: process.env.TOKENS_API_KEY });
const resp = await client.completions.create({
model: "deepseek/deepseek-v4.1-flash",
prompt: "A one-line definition of HTTP 429:",
max_tokens: 40,
temperature: 0,
});
console.log(resp.choices[0].text, resp.usage);If this returns an error, read Which models work. The Content-Type: application/json header matters: without it the gateway can't read the body and answers 400 invalid_request about a missing model field.
Request fields#
The body follows OpenAI's completions format. Checked against OpenAI's API reference in October 2026.
| Field | Type | Notes |
|---|---|---|
model | string | Required. A catalog id. List yours with GET /v1/models. |
prompt | string or array | Required in OpenAI's format. A string, or an array of strings. Token-id arrays are in the format too. |
max_tokens | integer | Upper bound on generated tokens. Set it; see the note on reservations below. |
temperature, top_p | number | Sampling. Change one or the other, as OpenAI recommends. |
n | integer | Number of completions per prompt. The gateway accepts 1 to 4 and returns 400 invalid_request otherwise. |
stop | string or array | Up to 4 stop sequences in OpenAI's format. |
stream | boolean | true returns Server-Sent Events. |
stream_options.include_usage | boolean | With stream: true, adds a final chunk with usage. |
suffix, echo, logprobs, best_of | various | OpenAI's format has them. Whether a model honors them is up to its provider. |
seed, presence_penalty, frequency_penalty, logit_bias, user | various | Passed through as sent. |
Parameters depend on the upstream model
Apart from model and n, the gateway doesn't validate or rewrite these fields. It passes them to the provider serving the model, which may ignore a field or reject the request. In OpenAI's own documentation, for example, suffix works only with one model. Check the model's page in the catalog and test.
The request body can be up to 10 MB. Larger bodies return 413 request_entity_too_large.
Example response#
Ids and numbers are illustrative.
{
"id": "cmpl-a1b2c3",
"object": "text_completion",
"created": 1790000000,
"model": "deepseek/deepseek-v4.1-flash",
"choices": [
{
"index": 0,
"text": " The server is rate limiting you; wait and retry.",
"finish_reason": "stop",
"logprobs": null
}
],
"usage": { "prompt_tokens": 12, "completion_tokens": 11, "total_tokens": 23 }
}The generated text is in choices[].text, not choices[].message.content as in chat completions. finish_reason is stop when the model ended on its own or hit a stop sequence, and length when it reached max_tokens. The body comes from the provider, so extra fields such as system_fingerprint vary by model.
Streaming#
Set "stream": true and the response is Server-Sent Events. Each event carries a partial choices[].text, and the stream ends with data: [DONE]:
data: {"id":"cmpl-a1b2c3","object":"text_completion","choices":[{"index":0,"text":" The server","finish_reason":null}]}
data: {"id":"cmpl-a1b2c3","object":"text_completion","choices":[{"index":0,"text":" is rate limiting you.","finish_reason":"stop"}]}
data: [DONE]For billing, the gateway asks the provider to include a usage chunk at the end of a streamed completion. That chunk has an empty choices array and a usage object, so write your stream parser to accept it. Disconnects, timeouts and proxy buffering behave as described in streaming.
Billing and limits#
- Metering. Usage is billed on the
prompt_tokensandcompletion_tokensthe provider reports, at the model's input and output prices. If a provider sends no usage, the gateway estimates from the request and response size. - Reservation. Before forwarding, the gateway reserves the worst-case cost of the request, using
max_tokensfor the output side. If you leave it out, the reservation assumes 8,192 output tokens, even though the provider's own default may be much smaller. A key close to its monthly spend cap or a balance near zero can refuse a request that omitsmax_tokenswhile a request with a small one passes. If your balance can't cover the requestedmax_tokens, the gateway may lower it to what the balance covers. You are charged for actual usage, not the reservation. - Rate limits. A completions request counts toward your per-minute limit and concurrency like any other inference request. See rate limits.
- Failed requests are not billed. A provider response with a status of 400 or above is returned without a charge.
Which models work#
The gateway forwards the call as is. It does not translate a completions request into a chat request, and it does not check whether the model supports plain completion. What you get back depends on the model's provider:
| Response | Likely cause |
|---|---|
| 200 with text | The provider serves /completions for this model. |
400 invalid_request | The provider refused the request, for example because the model is chat-only or a field isn't allowed. |
404 model_not_found | The provider has no completions route for this model. Also returned for an id not in the catalog. |
400 endpoint_not_supported_for_model | The model is served only through a provider that speaks the Anthropic Messages protocol. Nothing was sent. |
502 upstream_unreachable | The model has several providers and each one failed or refused. |
The upstream message is replaced with a generic one, so keep the x-tokens-request-id header if you need support to look at a failure. Full list in errors.
Move to chat completions#
If a completions call fails or you are writing new code, the same request as a chat completion is one change of shape:
# Before: /v1/completions
resp = client.completions.create(model=model, prompt="A one-line definition of HTTP 429:", max_tokens=40)
text = resp.choices[0].text
# After: /v1/chat/completions
resp = client.chat.completions.create(
model=model,
messages=[{"role": "user", "content": "A one-line definition of HTTP 429:"}],
max_tokens=40,
)
text = resp.choices[0].message.contentA chat model answers a question or instruction, where a completions model continues text, so reword prompts that relied on continuation (for example "The three causes are:" becomes "List the three causes."). Features that exist only on completions, such as suffix and echo, have no chat equivalent.
Related#
- Chat completions, the endpoint to use for new work.
- Responses and Messages for the other inference endpoints.
- Token counting for estimating prompt size and reading
usage.