Skip to content

Vision: sending images to models

Send images to a model on Chat Completions (image_url parts) or Messages (image blocks): URL and base64 forms, size limits, what the gateway translates, how image input is billed, and how to check a model accepts images.

On this page

Vision means sending an image to a model and asking about it: a screenshot of an error, a diagram, a photo of a whiteboard. On Tokens images travel inside the normal chat request, on POST https://tokens.bd/v1/chat/completions as image_url content parts and on POST https://tokens.bd/v1/messages as image content blocks. The gateway carries the image to the provider and bills what the provider reports. Whether the model can read the image is up to the model.

Tokens has no separate image endpoint. Image generation, file uploads and the Anthropic Files API are not served (see Models and usage). This page is about image input only.

Check that the model accepts images#

Not every model reads images. The model catalog at /models records the context window, prices and a description for each model, but it has no field that says "supports images". To find out:

  • Read the model's description on its /models page and the maker's own documentation for the model.
  • See the vision section of Choosing a model, which lists several models by what they accept.
  • Send one small test image (the examples below) and ask "What is in this image?". A model that reads images answers about the picture.

A text-only model does not always fail loudly. The provider decides: it can return a 400 (which reaches you as invalid_request, see errors), or it can drop the image and answer from the text alone. If the answer ignores the picture, treat that as "this model does not take images".

Chat Completions: image_url parts#

Put an array in the message content instead of a string. Each element is a text part or an image_url part. The url is either a public https URL or a base64 data URL (data:<media type>;base64,<data>).

json
{
  "role": "user",
  "content": [
    { "type": "text", "text": "What does this error screenshot say?" },
    {
      "type": "image_url",
      "image_url": { "url": "data:image/png;base64,<BASE64_DATA>", "detail": "auto" }
    }
  ]
}

detail is OpenAI's optional hint (low, high or auto; auto when omitted). Other providers may ignore it, and the gateway drops it when it has to translate the request for an Anthropic-style provider (see below).

curl https://tokens.bd/v1/chat/completions \
  -H "Authorization: Bearer $TOKENS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek/deepseek-v4.1-flash",
    "max_tokens": 300,
    "messages": [
      {
        "role": "user",
        "content": [
          {"type": "image_url", "image_url": {"url": "https://YOUR-HOST/path/to/image.png"}},
          {"type": "text", "text": "Describe this image in two sentences."}
        ]
      }
    ]
  }'

Replace https://YOUR-HOST/path/to/image.png with a link to an image you control. The media type in a data URL must match the file (image/png, image/jpeg, image/webp, image/gif).

Who fetches a URL image

When you send an https URL, Tokens does not download the image. It forwards the URL and the provider fetches it, so the link must be public and reachable from the provider. Some providers accept only base64 data URLs. If a URL fails and the same image works as base64, use base64.

Messages: image content blocks#

On /v1/messages an image is a content block with a source. Two source types work through Tokens:

json
{
  "role": "user",
  "content": [
    {
      "type": "image",
      "source": { "type": "base64", "media_type": "image/png", "data": "<BASE64_DATA>" }
    },
    { "type": "text", "text": "What does this error screenshot say?" }
  ]
}
json
{
  "type": "image",
  "source": { "type": "url", "url": "https://YOUR-HOST/path/to/image.png" }
}

data is the raw base64 string with no data: prefix. Anthropic's third source, {"type": "file", "file_id": "..."}, needs their Files API, which Tokens does not serve. Send the image itself.

curl https://tokens.bd/v1/messages \
  -H "x-api-key: $TOKENS_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{
    "model": "deepseek/deepseek-v4.1-flash",
    "max_tokens": 300,
    "messages": [
      {
        "role": "user",
        "content": [
          {"type": "image", "source": {"type": "url", "url": "https://YOUR-HOST/path/to/image.png"}},
          {"type": "text", "text": "Describe this image in two sentences."}
        ]
      }
    ]
  }'

Anthropic's guidance is to put images before the text that asks about them, and to label several images in the text ("Image 1:", "Image 2:") so you can refer to them. Earlier images in a conversation stay visible to the model, but you must resend them in every request, because the API is stateless.

Base64 from the command line#

A base64 image is too large to type into a curl -d '...' argument. Write the request to a file and send the file. On macOS and Linux, with jq installed:

bash
base64 screenshot.png | tr -d '\n' > screenshot.b64

jq -n --rawfile img screenshot.b64 '{
  model: "deepseek/deepseek-v4.1-flash",
  max_tokens: 300,
  messages: [{
    role: "user",
    content: [
      {type: "image_url", image_url: {url: ("data:image/png;base64," + $img)}},
      {type: "text", text: "Describe this image in two sentences."}
    ]
  }]
}' > request.json

curl https://tokens.bd/v1/chat/completions \
  -H "Authorization: Bearer $TOKENS_API_KEY" \
  -H "Content-Type: application/json" \
  -d @request.json

What the gateway does with images#

On a native request (the provider speaks the same format as your request) the gateway forwards the body and changes only the model field (plus a usage-reporting option on streamed Chat Completions). It does not decode, resize or inspect the image.

Some models are only available from a provider that speaks the other format. Then the gateway translates, and images are handled like this:

Your requestProvider speaksWhat happens to images
Messages (image block)OpenAI formatBase64 becomes a data URL and a URL stays a URL, as image_url parts of the user message. Images in assistant turns are dropped. Images inside a tool_result are not sent as images (see warning).
Chat Completions (image_url part)Anthropic formatA data URL becomes a base64 image block, any other URL becomes a URL image block. detail is dropped. Only images in user messages are carried.

Images inside tool results

On a Messages request that the gateway translates to the OpenAI format, a tool_result whose content includes an image block is flattened to text, and the image block is written into that text as JSON. The model does not see a picture and the base64 data counts as input tokens. If an agent returns screenshots from tools, use a model that is served in Anthropic format, or have the tool describe the image in text.

The same translation rules, and the other fields that are carried across, are described in Messages.

Size limits#

Two limits matter, and the smaller one wins.

Tokens: the whole request body can be at most 10 MB, otherwise the gateway answers 413 request_entity_too_large. Base64 makes data about a third larger than the file, so all images in one request together can be roughly 7 MB of image files at most, and less once the text and earlier conversation turns are counted. Every request resends the full conversation, including earlier images.

The provider: each provider has its own limits on image formats, dimensions, count and size per image. As checked in October 2026:

  • Anthropic: JPEG, PNG, GIF (first frame only) and WebP; up to 8000 x 8000 pixels per image; 10 MB per image (base64); up to 100 images per request on models with a 200K-token context window and 600 on others. Above 20 images in one request a stricter per-image pixel limit applies, and Anthropic suggests keeping each side to 2000 px. Source: Anthropic vision documentation.
  • OpenAI: PNG, JPEG, WebP and non-animated GIF; up to 1,500 images per request. Source: OpenAI images and vision guide. OpenAI's 512 MB payload limit is above Tokens' 10 MB, so the Tokens limit applies first.
  • Other makers publish their own limits. Check the maker's documentation for the model you use.

Resize large photos before sending. A phone photo is often 4000 pixels wide, and models downscale it anyway, so you pay for upload size and latency without gaining detail. Keep text in screenshots legible: very aggressive JPEG compression makes small text unreadable.

How image input is billed#

An image is converted to input tokens by the provider, and the provider reports the total in usage. Tokens bills that reported input count at the model's input price, like any other input. Roughly, providers count images in patches: Anthropic documents ceil(width / 28) x ceil(height / 28) tokens per image before downscaling to a cap (about 1,568 tokens on its standard tier, up to 4,784 on its high-resolution tier). OpenAI documents patch-based and tile-based counts that depend on the model and on detail. Other makers differ, so read usage from a test request to see the real cost for your model.

Admission works from a different number. Before forwarding a request the gateway reserves its worst-case cost against your balance, and it estimates the input side as one token per four characters of the request body. A large base64 image is a lot of characters, so the reservation for a request with a multi-megabyte image can be far above what the provider ends up charging. You are charged the real usage, not the reservation, but a low balance or a key close to its spend cap can be refused with insufficient_credits or monthly_spend_cap_exceeded for a request that would have cost very little. Resizing the image fixes both the cost and the refusal. See the reservation notes in Chat Completions.

Troubleshooting#

SymptomLikely causeFix
400 invalid_request mentioning images or contentThe model does not accept image input, or the media type is not one it supportsTry a model that takes images; check the data URL's media type matches the file.
The answer ignores the pictureThe provider dropped the image for a text-only modelSwitch to a model that reads images.
413 request_entity_too_largeBody over 10 MB, usually base64 images or a long history of imagesResize or compress, send fewer images, or drop old turns.
402 insufficient_credits with a small balanceThe input estimate counts the base64 charactersResize the image, or top up in billing.
URL image fails, base64 worksThe provider could not fetch the URL, or only accepts base64Use base64, or make the URL public.
Messages request with tool screenshots, model is blindtool_result images are flattened on translated modelsSee the warning above.

If the error does not say what is wrong, the upstream message is replaced with a generic one. Send the x-tokens-request-id to support, and see errors for the codes.

Was this page helpful?

Still stuck? Open a support ticket

Need help configuring your agent?

Test your connection with the connection tester, or create an API key.