# Vision: sending images to models

> Send images to a model on Chat Completions (image_url parts) or Messages (image blocks): URL and base64 forms, size limits, what the gateway translates, how image input is billed, and how to check a model accepts images.

Vision means sending an image to a model and asking about it: a screenshot of an error, a diagram, a photo of a whiteboard. On Tokens images travel inside the normal chat request, on `POST https://tokens.bd/v1/chat/completions` as `image_url` content parts and on `POST https://tokens.bd/v1/messages` as `image` content blocks. The gateway carries the image to the provider and bills what the provider reports. Whether the model can read the image is up to the model.

Tokens has no separate image endpoint. Image generation, file uploads and the Anthropic Files API are not served (see [Models and usage](/docs/models-and-usage)). This page is about image **input** only.

## Check that the model accepts images

Not every model reads images. The model catalog at [/models](/models) records the context window, prices and a description for each model, but it has no field that says "supports images". To find out:

- Read the model's description on its [/models](/models) page and the maker's own documentation for the model.
- See the vision section of [Choosing a model](/docs/choosing-a-model), which lists several models by what they accept.
- Send one small test image (the examples below) and ask "What is in this image?". A model that reads images answers about the picture.

A text-only model does not always fail loudly. The provider decides: it can return a 400 (which reaches you as `invalid_request`, see [errors](/docs/errors)), or it can drop the image and answer from the text alone. If the answer ignores the picture, treat that as "this model does not take images".

## Chat Completions: image_url parts

Put an array in the message `content` instead of a string. Each element is a `text` part or an `image_url` part. The `url` is either a public `https` URL or a base64 data URL (`data:<media type>;base64,<data>`).

```json
{
  "role": "user",
  "content": [
    { "type": "text", "text": "What does this error screenshot say?" },
    {
      "type": "image_url",
      "image_url": { "url": "data:image/png;base64,<BASE64_DATA>", "detail": "auto" }
    }
  ]
}
```

`detail` is OpenAI's optional hint (`low`, `high` or `auto`; `auto` when omitted). Other providers may ignore it, and the gateway drops it when it has to translate the request for an Anthropic-style provider (see below).

:::code-tabs

```bash title="cURL (public URL)"
curl https://tokens.bd/v1/chat/completions \
  -H "Authorization: Bearer $TOKENS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek/deepseek-v4.1-flash",
    "max_tokens": 300,
    "messages": [
      {
        "role": "user",
        "content": [
          {"type": "image_url", "image_url": {"url": "https://YOUR-HOST/path/to/image.png"}},
          {"type": "text", "text": "Describe this image in two sentences."}
        ]
      }
    ]
  }'
```

```python title="Python (local file)"
import base64
import os
from openai import OpenAI

client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"])

with open("screenshot.png", "rb") as f:
    b64 = base64.b64encode(f.read()).decode("ascii")

resp = client.chat.completions.create(
    model="deepseek/deepseek-v4.1-flash",
    max_tokens=300,
    messages=[
        {
            "role": "user",
            "content": [
                {"type": "image_url", "image_url": {"url": f"data:image/png;base64,{b64}"}},
                {"type": "text", "text": "Describe this image in two sentences."},
            ],
        }
    ],
)
print(resp.choices[0].message.content)
print(resp.usage)
```

```typescript title="Node.js (local file)"
import fs from "node:fs";
import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://tokens.bd/v1", apiKey: process.env.TOKENS_API_KEY });

const b64 = fs.readFileSync("screenshot.png").toString("base64");

const resp = await client.chat.completions.create({
  model: "deepseek/deepseek-v4.1-flash",
  max_tokens: 300,
  messages: [
    {
      role: "user",
      content: [
        { type: "image_url", image_url: { url: `data:image/png;base64,${b64}` } },
        { type: "text", text: "Describe this image in two sentences." },
      ],
    },
  ],
});
console.log(resp.choices[0].message.content, resp.usage);
```

:::

Replace `https://YOUR-HOST/path/to/image.png` with a link to an image you control. The `media type` in a data URL must match the file (`image/png`, `image/jpeg`, `image/webp`, `image/gif`).

:::note[Who fetches a URL image]
When you send an `https` URL, Tokens does not download the image. It forwards the URL and the provider fetches it, so the link must be public and reachable from the provider. Some providers accept only base64 data URLs. If a URL fails and the same image works as base64, use base64.
:::

## Messages: image content blocks

On `/v1/messages` an image is a content block with a `source`. Two source types work through Tokens:

```json
{
  "role": "user",
  "content": [
    {
      "type": "image",
      "source": { "type": "base64", "media_type": "image/png", "data": "<BASE64_DATA>" }
    },
    { "type": "text", "text": "What does this error screenshot say?" }
  ]
}
```

```json
{
  "type": "image",
  "source": { "type": "url", "url": "https://YOUR-HOST/path/to/image.png" }
}
```

`data` is the raw base64 string with no `data:` prefix. Anthropic's third source, `{"type": "file", "file_id": "..."}`, needs their Files API, which Tokens does not serve. Send the image itself.

:::code-tabs

```bash title="cURL (public URL)"
curl https://tokens.bd/v1/messages \
  -H "x-api-key: $TOKENS_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{
    "model": "deepseek/deepseek-v4.1-flash",
    "max_tokens": 300,
    "messages": [
      {
        "role": "user",
        "content": [
          {"type": "image", "source": {"type": "url", "url": "https://YOUR-HOST/path/to/image.png"}},
          {"type": "text", "text": "Describe this image in two sentences."}
        ]
      }
    ]
  }'
```

```python title="Python (local file)"
import base64
import os
import anthropic

client = anthropic.Anthropic(base_url="https://tokens.bd", api_key=os.environ["TOKENS_API_KEY"])

with open("screenshot.png", "rb") as f:
    b64 = base64.standard_b64encode(f.read()).decode("ascii")

message = client.messages.create(
    model="deepseek/deepseek-v4.1-flash",
    max_tokens=300,
    messages=[
        {
            "role": "user",
            "content": [
                {"type": "image", "source": {"type": "base64", "media_type": "image/png", "data": b64}},
                {"type": "text", "text": "Describe this image in two sentences."},
            ],
        }
    ],
)
print(message.content[0].text)
print(message.usage)
```

```typescript title="Node.js (local file)"
import fs from "node:fs";
import Anthropic from "@anthropic-ai/sdk";

const client = new Anthropic({ baseURL: "https://tokens.bd", apiKey: process.env.TOKENS_API_KEY });

const b64 = fs.readFileSync("screenshot.png").toString("base64");

const message = await client.messages.create({
  model: "deepseek/deepseek-v4.1-flash",
  max_tokens: 300,
  messages: [
    {
      role: "user",
      content: [
        { type: "image", source: { type: "base64", media_type: "image/png", data: b64 } },
        { type: "text", text: "Describe this image in two sentences." },
      ],
    },
  ],
});
console.log(message.content[0], message.usage);
```

:::

Anthropic's guidance is to put images before the text that asks about them, and to label several images in the text ("Image 1:", "Image 2:") so you can refer to them. Earlier images in a conversation stay visible to the model, but you must resend them in every request, because the API is stateless.

## Base64 from the command line

A base64 image is too large to type into a `curl -d '...'` argument. Write the request to a file and send the file. On macOS and Linux, with `jq` installed:

```bash
base64 screenshot.png | tr -d '\n' > screenshot.b64

jq -n --rawfile img screenshot.b64 '{
  model: "deepseek/deepseek-v4.1-flash",
  max_tokens: 300,
  messages: [{
    role: "user",
    content: [
      {type: "image_url", image_url: {url: ("data:image/png;base64," + $img)}},
      {type: "text", text: "Describe this image in two sentences."}
    ]
  }]
}' > request.json

curl https://tokens.bd/v1/chat/completions \
  -H "Authorization: Bearer $TOKENS_API_KEY" \
  -H "Content-Type: application/json" \
  -d @request.json
```

## What the gateway does with images

On a native request (the provider speaks the same format as your request) the gateway forwards the body and changes only the `model` field (plus a usage-reporting option on streamed Chat Completions). It does not decode, resize or inspect the image.

Some models are only available from a provider that speaks the other format. Then the gateway translates, and images are handled like this:

| Your request                                    | Provider speaks | What happens to images                                                                                                                                                                                  |
| ----------------------------------------------- | --------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Messages (`image` block)                        | OpenAI format   | Base64 becomes a data URL and a URL stays a URL, as `image_url` parts of the user message. Images in `assistant` turns are dropped. Images inside a `tool_result` are not sent as images (see warning). |
| Chat Completions (`image_url` part)             | Anthropic format | A data URL becomes a base64 `image` block, any other URL becomes a URL `image` block. `detail` is dropped. Only images in `user` messages are carried.                                                 |

:::warning[Images inside tool results]
On a Messages request that the gateway translates to the OpenAI format, a `tool_result` whose content includes an `image` block is flattened to text, and the image block is written into that text as JSON. The model does not see a picture and the base64 data counts as input tokens. If an agent returns screenshots from tools, use a model that is served in Anthropic format, or have the tool describe the image in text.
:::

The same translation rules, and the other fields that are carried across, are described in [Messages](/docs/messages).

## Size limits

Two limits matter, and the smaller one wins.

**Tokens:** the whole request body can be at most 10 MB, otherwise the gateway answers `413 request_entity_too_large`. Base64 makes data about a third larger than the file, so all images in one request together can be roughly 7 MB of image files at most, and less once the text and earlier conversation turns are counted. Every request resends the full conversation, including earlier images.

**The provider:** each provider has its own limits on image formats, dimensions, count and size per image. As checked in October 2026:

- Anthropic: JPEG, PNG, GIF (first frame only) and WebP; up to 8000 x 8000 pixels per image; 10 MB per image (base64); up to 100 images per request on models with a 200K-token context window and 600 on others. Above 20 images in one request a stricter per-image pixel limit applies, and Anthropic suggests keeping each side to 2000 px. Source: [Anthropic vision documentation](https://platform.claude.com/docs/en/build-with-claude/vision).
- OpenAI: PNG, JPEG, WebP and non-animated GIF; up to 1,500 images per request. Source: [OpenAI images and vision guide](https://developers.openai.com/api/docs/guides/images-vision). OpenAI's 512 MB payload limit is above Tokens' 10 MB, so the Tokens limit applies first.
- Other makers publish their own limits. Check the maker's documentation for the model you use.

Resize large photos before sending. A phone photo is often 4000 pixels wide, and models downscale it anyway, so you pay for upload size and latency without gaining detail. Keep text in screenshots legible: very aggressive JPEG compression makes small text unreadable.

## How image input is billed

An image is converted to input tokens by the provider, and the provider reports the total in `usage`. Tokens bills that reported input count at the model's input price, like any other input. Roughly, providers count images in patches: Anthropic documents `ceil(width / 28) x ceil(height / 28)` tokens per image before downscaling to a cap (about 1,568 tokens on its standard tier, up to 4,784 on its high-resolution tier). OpenAI documents patch-based and tile-based counts that depend on the model and on `detail`. Other makers differ, so read `usage` from a test request to see the real cost for your model.

Admission works from a different number. Before forwarding a request the gateway reserves its worst-case cost against your balance, and it estimates the input side as one token per four characters of the request body. A large base64 image is a lot of characters, so the reservation for a request with a multi-megabyte image can be far above what the provider ends up charging. You are charged the real usage, not the reservation, but a low balance or a key close to its spend cap can be refused with `insufficient_credits` or `monthly_spend_cap_exceeded` for a request that would have cost very little. Resizing the image fixes both the cost and the refusal. See the reservation notes in [Chat Completions](/docs/chat-completions).

## Troubleshooting

| Symptom                                               | Likely cause                                                           | Fix                                                                                                               |
| ----------------------------------------------------- | ---------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------- |
| `400 invalid_request` mentioning images or content    | The model does not accept image input, or the media type is not one it supports | Try a model that takes images; check the data URL's media type matches the file.                         |
| The answer ignores the picture                        | The provider dropped the image for a text-only model                   | Switch to a model that reads images.                                                                              |
| `413 request_entity_too_large`                        | Body over 10 MB, usually base64 images or a long history of images     | Resize or compress, send fewer images, or drop old turns.                                                         |
| `402 insufficient_credits` with a small balance       | The input estimate counts the base64 characters                        | Resize the image, or top up in [billing](/dashboard/billing).                                                      |
| URL image fails, base64 works                         | The provider could not fetch the URL, or only accepts base64           | Use base64, or make the URL public.                                                                               |
| Messages request with tool screenshots, model is blind | `tool_result` images are flattened on translated models                | See the warning above.                                                                                            |

If the error does not say what is wrong, the upstream message is replaced with a generic one. Send the `x-tokens-request-id` to [support](/docs/support), and see [errors](/docs/errors) for the codes.

---
Page: https://tokens.bd/docs/vision
