Skip to content

Embeddings

POST /v1/embeddings turns text into vectors for search, retrieval and clustering. How to find an embedding model in the catalog, the request and response, how it is billed, and the errors you can hit.

On this page

POST https://tokens.bd/v1/embeddings is the OpenAI-compatible embeddings endpoint. You send text and get back one vector per input, which you can store and compare for semantic search, retrieval for RAG, deduplication or clustering. The gateway checks your key, plan and limits, then forwards the body to the upstream provider for the model you named.

This endpoint only works with embedding models. A chat model such as deepseek/deepseek-v4.1-flash is not one, and sending it here fails. Read the next section first.

Find an embedding model#

Tokens does not add a capability flag to GET /v1/models, so the list has ids and nothing else. To find an embedding model:

  1. Open the model catalog and search for embed. The search matches a model's name, id and provider.
  2. Open the model's page to check its context window and its input price per million tokens.
  3. Confirm your key can call it: the id must appear in GET /v1/models for that key (why a model can be missing).

If no model in the catalog is an embedding model, none is available to your account yet and /v1/embeddings has nothing to serve. Keep your current embedding provider until one is listed. This page does not name a model because the catalog changes; the examples below read the id from an environment variable instead:

bash
export EMBEDDING_MODEL="the-id-from-the-catalog"

Create embeddings#

curl https://tokens.bd/v1/embeddings \
  -H "Authorization: Bearer $TOKENS_API_KEY" \
  -H "Content-Type: application/json" \
  -d @- <<EOF
{
  "model": "$EMBEDDING_MODEL",
  "input": ["How do I reset my password?", "Where can I change my billing email?"],
  "encoding_format": "float"
}
EOF

The Content-Type: application/json header matters. Without it the gateway can't read the body and answers 400 invalid_request saying the request must specify a model field, even though it does.

Request fields#

The body follows OpenAI's embeddings format. Checked against OpenAI's API reference in October 2026.

FieldTypeNotes
modelstringRequired. The id of an embedding model, exactly as the catalog shows it.
inputstring or arrayRequired. One string, or an array of strings to embed in one request. OpenAI's format also allows token ids.
encoding_formatstring"float" (default in the raw API) or "base64". Support is up to the model's provider.
dimensionsintegerShorter output vectors. In OpenAI's own API only some models accept it; whether yours does is decided by its provider.
userstringAn end-user identifier, passed to the provider.

Parameters depend on the upstream model

The gateway reads model and nothing else from this body. The other fields go to the provider unchanged, so the limits that matter (maximum tokens per input, how many inputs per request, whether dimensions is allowed, the vector length) are the provider's, and they differ between models. For OpenAI's own embedding models, OpenAI documents 8,192 tokens per input, at most 2,048 items in an input array and 300,000 tokens across one request. Do not assume the same numbers for another model.

The request body can be up to 10 MB in total. Larger bodies return 413 request_entity_too_large.

The Python SDK asks for base64 by default#

If you leave encoding_format out, the OpenAI Python SDK sends "base64" and decodes the result itself (checked in the SDK source, October 2026). That works only when the model's provider supports base64. If a call fails or returns odd vectors, set encoding_format="float" as the examples above do. Setting it explicitly in the Node.js SDK costs nothing either.

Example response#

Vectors are shortened here. A real one has hundreds or thousands of numbers, so the response body is large compared with the request.

json
{
  "object": "list",
  "data": [
    { "object": "embedding", "index": 0, "embedding": [0.0123, -0.0456, 0.0789] },
    { "object": "embedding", "index": 1, "embedding": [0.0311, -0.0127, 0.0644] }
  ],
  "model": "the-id-from-the-catalog",
  "usage": { "prompt_tokens": 14, "total_tokens": 14 }
}

data has one entry per input, in the same order; index is the position of the input. The body comes from the provider, so extra fields can appear and the vector length depends on the model. There is no completion_tokens because an embedding request produces no text.

Compare two vectors#

Vectors from the same model can be compared with cosine similarity. A small example in Python:

python
import math

def cosine(a, b):
    dot = sum(x * y for x, y in zip(a, b))
    return dot / (math.sqrt(sum(x * x for x in a)) * math.sqrt(sum(y * y for y in b)))

print(cosine(vectors[0], vectors[1]))

Never compare or index together vectors from different models, or from the same model with different dimensions. Store the model id and dimensions next to each vector so you know when to re-embed.

Billing and limits#

  • Metering. Embeddings are billed on the usage.prompt_tokens the provider reports, at the model's input price per million tokens. There is no output side. Per-model prices are in the model catalog and pricing. If a provider sends no usage, the gateway estimates from the request and response size.
  • No max_tokens. Unlike chat, there is no output reservation. Admission reserves the worst-case cost from the request size only, so a request is refused for funds only when your balance can't cover that estimate. You pay for actual usage.
  • Rate limits. Every embeddings request counts as one request toward your per-minute limit and holds a concurrency slot while it runs, whatever its size. Put many inputs into one input array instead of sending one request per text, within the provider's limit for the model. See rate limits.
  • Failed requests are not billed. A response with a status of 400 or above from the provider is returned to you without a charge.
  • Not streamed. Embeddings have no stream mode.

Using a chat model on this endpoint#

Sending a chat model here is the most common mistake. What you see depends on the model's provider, not on a fixed Tokens rule:

ResponseCause
400 invalid_request, "rejected by the upstream provider"The provider took the request and refused it because the model can't produce embeddings.
404 model_not_foundThe provider doesn't have an embeddings route for that model. Also returned for an id that isn't in the catalog.
400 endpoint_not_supported_for_modelThe model is served only through a provider that speaks the Anthropic Messages protocol, which has no embeddings. Nothing was sent.
502 upstream_unreachableThe model has several providers and each one failed or refused.

In every case switch to an embedding model from the catalog. The upstream message is replaced with a generic one, so use the x-tokens-request-id header when you contact support. The full list is in errors.

Errors#

Gateway errors use the OpenAI error shape with a code you can branch on, and every response carries an x-tokens-request-id header. The ones you meet most often here:

StatusCodeWhat to do
400invalid_requestCheck model and input, and that you sent Content-Type: application/json.
402insufficient_credits, no_fundingTop up or renew in billing.
403model_not_allowed_on_keyThe key has an allow-list that excludes this model. Use another key.
404model_not_foundCheck the id with GET /v1/models, or the model isn't an embedding model.
413request_entity_too_largeSend fewer inputs per request.
429rate_limited, concurrency_limitWait for Retry-After; batch inputs into fewer requests.

For a bulk indexing job, add retries with backoff as shown in rate limits, and keep the number of parallel requests under your plan's concurrency limit.

Was this page helpful?

Still stuck? Open a support ticket

Need help configuring your agent?

Test your connection with the connection tester, or create an API key.