# Embeddings

> POST /v1/embeddings turns text into vectors for search, retrieval and clustering. How to find an embedding model in the catalog, the request and response, how it is billed, and the errors you can hit.

`POST https://tokens.bd/v1/embeddings` is the OpenAI-compatible embeddings endpoint. You send text and get back one vector per input, which you can store and compare for semantic search, retrieval for RAG, deduplication or clustering. The gateway checks your key, plan and limits, then forwards the body to the upstream provider for the model you named.

This endpoint only works with **embedding models**. A chat model such as `deepseek/deepseek-v4.1-flash` is not one, and sending it here fails. Read the next section first.

## Find an embedding model

Tokens does not add a capability flag to `GET /v1/models`, so the list has ids and nothing else. To find an embedding model:

1. Open [the model catalog](/models) and search for `embed`. The search matches a model's name, id and provider.
2. Open the model's page to check its context window and its input price per million tokens.
3. Confirm your key can call it: the id must appear in `GET /v1/models` for that key ([why a model can be missing](/docs/models-and-usage)).

If no model in the catalog is an embedding model, none is available to your account yet and `/v1/embeddings` has nothing to serve. Keep your current embedding provider until one is listed. This page does not name a model because the catalog changes; the examples below read the id from an environment variable instead:

```bash
export EMBEDDING_MODEL="the-id-from-the-catalog"
```

## Create embeddings

:::code-tabs

```bash title="cURL"
curl https://tokens.bd/v1/embeddings \
  -H "Authorization: Bearer $TOKENS_API_KEY" \
  -H "Content-Type: application/json" \
  -d @- <<EOF
{
  "model": "$EMBEDDING_MODEL",
  "input": ["How do I reset my password?", "Where can I change my billing email?"],
  "encoding_format": "float"
}
EOF
```

```python title="Python"
import os
from openai import OpenAI

client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"])

resp = client.embeddings.create(
    model=os.environ["EMBEDDING_MODEL"],
    input=["How do I reset my password?", "Where can I change my billing email?"],
    encoding_format="float",
)

vectors = [item.embedding for item in resp.data]
print(len(vectors), "vectors of", len(vectors[0]), "numbers")
print(resp.usage)
```

```typescript title="Node.js"
import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://tokens.bd/v1", apiKey: process.env.TOKENS_API_KEY });

const resp = await client.embeddings.create({
  model: process.env.EMBEDDING_MODEL!,
  input: ["How do I reset my password?", "Where can I change my billing email?"],
  encoding_format: "float",
});

const vectors = resp.data.map((item) => item.embedding);
console.log(vectors.length, "vectors of", vectors[0].length, "numbers");
console.log(resp.usage);
```

:::

The `Content-Type: application/json` header matters. Without it the gateway can't read the body and answers 400 `invalid_request` saying the request must specify a `model` field, even though it does.

## Request fields

The body follows OpenAI's embeddings format. Checked against OpenAI's API reference in October 2026.

| Field             | Type             | Notes                                                                                                                      |
| ----------------- | ---------------- | -------------------------------------------------------------------------------------------------------------------------- |
| `model`           | string           | Required. The id of an embedding model, exactly as the catalog shows it.                                                   |
| `input`           | string or array  | Required. One string, or an array of strings to embed in one request. OpenAI's format also allows token ids.               |
| `encoding_format` | string           | `"float"` (default in the raw API) or `"base64"`. Support is up to the model's provider.                                    |
| `dimensions`      | integer          | Shorter output vectors. In OpenAI's own API only some models accept it; whether yours does is decided by its provider.      |
| `user`            | string           | An end-user identifier, passed to the provider.                                                                            |

:::note[Parameters depend on the upstream model]
The gateway reads `model` and nothing else from this body. The other fields go to the provider unchanged, so the limits that matter (maximum tokens per input, how many inputs per request, whether `dimensions` is allowed, the vector length) are the provider's, and they differ between models. For OpenAI's own embedding models, OpenAI documents 8,192 tokens per input, at most 2,048 items in an input array and 300,000 tokens across one request. Do not assume the same numbers for another model.
:::

The request body can be up to 10 MB in total. Larger bodies return 413 `request_entity_too_large`.

### The Python SDK asks for base64 by default

If you leave `encoding_format` out, the OpenAI Python SDK sends `"base64"` and decodes the result itself (checked in the SDK source, October 2026). That works only when the model's provider supports base64. If a call fails or returns odd vectors, set `encoding_format="float"` as the examples above do. Setting it explicitly in the Node.js SDK costs nothing either.

## Example response

Vectors are shortened here. A real one has hundreds or thousands of numbers, so the response body is large compared with the request.

```json
{
  "object": "list",
  "data": [
    { "object": "embedding", "index": 0, "embedding": [0.0123, -0.0456, 0.0789] },
    { "object": "embedding", "index": 1, "embedding": [0.0311, -0.0127, 0.0644] }
  ],
  "model": "the-id-from-the-catalog",
  "usage": { "prompt_tokens": 14, "total_tokens": 14 }
}
```

`data` has one entry per input, in the same order; `index` is the position of the input. The body comes from the provider, so extra fields can appear and the vector length depends on the model. There is no `completion_tokens` because an embedding request produces no text.

## Compare two vectors

Vectors from the same model can be compared with cosine similarity. A small example in Python:

```python
import math

def cosine(a, b):
    dot = sum(x * y for x, y in zip(a, b))
    return dot / (math.sqrt(sum(x * x for x in a)) * math.sqrt(sum(y * y for y in b)))

print(cosine(vectors[0], vectors[1]))
```

Never compare or index together vectors from different models, or from the same model with different `dimensions`. Store the model id and `dimensions` next to each vector so you know when to re-embed.

## Billing and limits

- **Metering.** Embeddings are billed on the `usage.prompt_tokens` the provider reports, at the model's input price per million tokens. There is no output side. Per-model prices are in [the model catalog](/models) and [pricing](/pricing). If a provider sends no usage, the gateway estimates from the request and response size.
- **No `max_tokens`.** Unlike chat, there is no output reservation. Admission reserves the worst-case cost from the request size only, so a request is refused for funds only when your balance can't cover that estimate. You pay for actual usage.
- **Rate limits.** Every embeddings request counts as one request toward your per-minute limit and holds a concurrency slot while it runs, whatever its size. Put many inputs into one `input` array instead of sending one request per text, within the provider's limit for the model. See [rate limits](/docs/rate-limits).
- **Failed requests are not billed.** A response with a status of 400 or above from the provider is returned to you without a charge.
- **Not streamed.** Embeddings have no `stream` mode.

## Using a chat model on this endpoint

Sending a chat model here is the most common mistake. What you see depends on the model's provider, not on a fixed Tokens rule:

| Response                                               | Cause                                                                                                                            |
| ------------------------------------------------------ | -------------------------------------------------------------------------------------------------------------------------------- |
| 400 `invalid_request`, "rejected by the upstream provider" | The provider took the request and refused it because the model can't produce embeddings.                                    |
| 404 `model_not_found`                                  | The provider doesn't have an embeddings route for that model. Also returned for an id that isn't in the catalog.                  |
| 400 `endpoint_not_supported_for_model`                 | The model is served only through a provider that speaks the Anthropic Messages protocol, which has no embeddings. Nothing was sent. |
| 502 `upstream_unreachable`                             | The model has several providers and each one failed or refused.                                                                   |

In every case switch to an embedding model from the catalog. The upstream message is replaced with a generic one, so use the `x-tokens-request-id` header when you contact support. The full list is in [errors](/docs/errors).

## Errors

Gateway errors use the OpenAI error shape with a `code` you can branch on, and every response carries an `x-tokens-request-id` header. The ones you meet most often here:

| Status | Code                                   | What to do                                                                       |
| ------ | -------------------------------------- | -------------------------------------------------------------------------------- |
| 400    | `invalid_request`                      | Check `model` and `input`, and that you sent `Content-Type: application/json`.   |
| 402    | `insufficient_credits`, `no_funding`   | Top up or renew in [billing](/dashboard/billing).                                |
| 403    | `model_not_allowed_on_key`             | The key has an allow-list that excludes this model. Use another key.             |
| 404    | `model_not_found`                      | Check the id with `GET /v1/models`, or the model isn't an embedding model.       |
| 413    | `request_entity_too_large`             | Send fewer inputs per request.                                                   |
| 429    | `rate_limited`, `concurrency_limit`    | Wait for `Retry-After`; batch inputs into fewer requests.                        |

For a bulk indexing job, add retries with backoff as shown in [rate limits](/docs/rate-limits), and keep the number of parallel requests under your plan's concurrency limit.

## Related

- [Chat completions](/docs/chat-completions) for text generation.
- [Token counting](/docs/token-counting) for estimating input size before you embed a large corpus.
- [Models and usage](/docs/models-and-usage) for `GET /v1/models` and `GET /v1/tokens/usage`.

---
Page: https://tokens.bd/docs/embeddings
