# Working with Bengali text

> How Bengali and other non-Latin text behaves through the API: why it takes more tokens, how to measure it from the usage field, what that means for max_tokens, context and cost, prompting tips, UTF-8 in streams, and a script to compare models on your own text.

Bengali works through Tokens like any other text: you send it as UTF-8 in the `messages` and read the answer back. Three things differ from English. The same content usually uses more tokens, so it costs more and fills the context window faster. Some bugs only show up with multi-byte characters, especially in streams. And you should check output quality on your own text before you pick a model. This page covers each, and ends with a script that compares models on a Bengali prompt.

The same points apply, to different degrees, to Hindi, Arabic, Thai, Chinese and other non-Latin scripts.

## Why Bengali takes more tokens

Models do not read characters. A tokenizer first splits the text into pieces called tokens, and you pay per token. A tokenizer is built from a large body of text, and a script that is rare in that text tends to be split into more, smaller pieces. How much depends on the model's tokenizer, and the maker can change it between versions. Anthropic, for example, documents that newer Claude models use a different tokenizer from older ones and that the same text produces a different number of tokens, and tells you to recount prompts against the model you plan to use ([Anthropic: token counting](https://platform.claude.com/docs/en/build-with-claude/token-counting), checked October 2026).

Bengali text is also built from many small Unicode parts. You can see this in Python:

```python
text = "বাংলা"
print(len(text))                  # 5 Unicode code points
print(len(text.encode("utf-8")))  # 15 bytes: 3 bytes for each Bengali character
print(len("ক্ষ"))                  # 3: ক + hasanta (virama) + ষ make one conjunct
```

Every Bengali letter takes three bytes in UTF-8, and a conjunct is several code points. A tokenizer that works on bytes starts from those pieces. How far it merges them back into larger tokens depends on whether Bengali was well represented in its training text, and that differs from model to model.

Do not rely on a ratio you read somewhere. Measure it on your own text and your own model, as shown next.

## Measure it with a real response

Every response reports its token counts in `usage`. Send an English text and a Bengali text with the same meaning, and compare `prompt_tokens`. Set `max_tokens` low so the call costs almost nothing: you are billed for the input you send, and for at most a few output tokens.

```python title="measure_tokens.py"
import os
from openai import OpenAI

client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"])

ENGLISH = (
    "Our team launched a new payment page last week. In the first two days, "
    "some users got an error when paying by card. The cause was an old library, "
    "and we have updated it."
)
BENGALI = (
    "আমাদের দল গত সপ্তাহে নতুন পেমেন্ট পেজ চালু করেছে। প্রথম দুই দিনে কয়েকজন ব্যবহারকারী "
    "কার্ড দিয়ে টাকা দিতে গিয়ে ত্রুটি পেয়েছেন। সমস্যাটি একটি পুরনো লাইব্রেরির কারণে হয়েছিল, "
    "এবং আমরা সেটি আপডেট করেছি।"
)

def input_tokens(model: str, text: str) -> int:
    resp = client.chat.completions.create(
        model=model,
        messages=[{"role": "user", "content": text}],
        max_tokens=16,  # the input is what we measure; keep the output tiny
    )
    return resp.usage.prompt_tokens

for model in ["deepseek/deepseek-v4.1-flash"]:  # add more ids from /models
    en = input_tokens(model, ENGLISH)
    bn = input_tokens(model, BENGALI)
    print(f"{model}: English {en} tokens, Bengali {bn} tokens, ratio {bn / en:.2f}")
```

Notes on reading the result:

- `prompt_tokens` includes a few tokens the model's chat format adds around your message. For very short texts that overhead distorts the ratio, so measure paragraphs of realistic length.
- Run it for **each model you might use**. Different models can give very different counts for the same text.
- Use your real prompts. A system prompt, tool definitions and pasted code are mostly English or code, and only part of your input is Bengali.
- Some models that think before answering may reject a very small `max_tokens`. If you get an error, raise it a little.

You can also count tokens before you send; see [token counting](/docs/token-counting).

## What it means for `max_tokens`, context and cost

- **`max_tokens`** counts output tokens. A Bengali answer of the same length needs more of them, so a limit that is comfortable in English can cut a Bengali answer short. Check `finish_reason`: `length` means your limit stopped it. Raise `max_tokens` for Bengali answers, then confirm the result by measuring.
- **Context window.** The window is measured in tokens, not characters. A conversation in Bengali reaches the limit sooner than the same conversation in English, so trim or summarize history earlier. Limits per model are on [/models](/models).
- **Cost.** You pay per token at the model's price, so the cost of a Bengali request is its token count times the price. Compare models by the cost of your Bengali task, not by price per token alone: a model with a lower price and a worse tokenizer for Bengali can cost more per request. See [cutting your spend](/docs/cutting-costs).
- **Speed.** More tokens to write means longer to finish, even at the same tokens per second.

## Prompting for Bengali

These are habits that make output predictable. Test each on your own task.

- **Say the output language explicitly.** A model answers in the language it considers likely, which may follow your instructions, the user message, or the system prompt. Put it in the system prompt, for example `Always answer in Bengali (Bangla script).` If you want English, say so just as plainly.
- **Name the script.** "Bengali" can come back in Bangla script or in Latin letters (romanized Bengali). Ask for Bangla script if that is what you need.
- **Keep code and identifiers in English.** Variable names, function names, file paths, JSON keys and error messages are best left in English. Write instructions like `Explain in Bengali, but keep code, identifiers and JSON keys unchanged.`
- **Machine-readable output.** If your code parses the answer, ask for a fixed format and say which digits to use. Bengali digits (০১২৩) and ASCII digits (0123) are different characters, and `int("১২")` works in Python while a regular expression for `[0-9]` will not match it. Say `Use ASCII digits` if you need them.
- **Give an example in the format you want.** One short Bengali example in the prompt shows tone and script better than a description does.
- **Mind mixed-script text.** Bengali with English technical terms is normal, but similar-looking characters can cause bugs. The Bengali danda `।` (U+0964) is not a pipe `|` or a Latin `I`, and the Bengali digit zero `০` (U+09E6) is not the Latin `0`. Search, comparison and parsing treat them as different characters.
- **Normalize Unicode before you compare or store.** The same Bengali text can be encoded in more than one way. For example `য়` can be one code point (U+09DF) or two (U+09AF followed by U+09BC), and the vowel sign in `কো` can be one code point (U+09CB) or two (U+09C7 followed by U+09BE). Normalize with NFC before you compare, search or deduplicate:

```python
import unicodedata

def clean(text: str) -> str:
    return unicodedata.normalize("NFC", text)

assert clean("য়") == "য়"          # য় composed form becomes two code points
assert clean("কো") == "কো"  # কো, two vowel parts become one sign
```

Normalizing the text you send also keeps your prompts consistent, which helps if you test or cache them. It does not change how the model writes its answer.

## Streaming and UTF-8 in clients

With `stream: true` the answer arrives in chunks. If you use an official SDK, it decodes the stream for you and you can skip this section. If you read the response bytes yourself, remember that one Bengali character is three bytes, and a network chunk can end in the middle of one. Decoding each chunk on its own then gives a replacement character (`�`) or an error.

Use an incremental decoder that keeps the unfinished bytes until the rest arrives.

:::code-tabs

```python title="Python"
import codecs

decoder = codecs.getincrementaldecoder("utf-8")()

def feed(chunk: bytes) -> str:
    return decoder.decode(chunk)       # returns only complete characters

data = "বা".encode("utf-8")            # 6 bytes
print(repr(data[:2].decode("utf-8", errors="replace")))  # '�' when split naively
print(repr(feed(data[:2]) + feed(data[2:])))             # 'বা' with the incremental decoder
```

```javascript title="Node.js / browser"
const decoder = new TextDecoder("utf-8");

function feed(chunk) {
  // { stream: true } keeps an unfinished character until the next chunk
  return decoder.decode(chunk, { stream: true });
}
```

:::

Related points:

- Decode first, then split into lines. Splitting raw bytes on newlines is safe for UTF-8 (the newline byte never appears inside a multi-byte character), but splitting a decoded string at a byte offset is not.
- When you cut text to a length, cut on characters, not on bytes. A byte-based cut can end in the middle of a character.
- Send requests with `Content-Type: application/json`, and make sure your HTTP library encodes the body as UTF-8. Do not write the JSON by escaping each character by hand.
- Read the server-sent events format in [streaming](/docs/streaming). To get token counts at the end of a stream, add `"stream_options": {"include_usage": true}` to the request.
- If you display text in a terminal or on Windows, the display font and console code page can show boxes or `?` for text that is correct in memory. Check by writing the string to a UTF-8 file.

## Check which models handle Bengali well

There is no reliable shortcut. A maker's claim that a model supports a language does not tell you how it does on your task, and a score on a public test does not tell you how it does on your text. Do not choose from a ranking. Test:

1. **Collect 20 to 50 real prompts.** Use the kind of text your product handles: support messages, summaries, code explanations, form data.
2. **Run the same prompts on several models** with the script below, and keep the answers.
3. **Read the answers yourself, or have a Bengali speaker read them.** Look at correctness first, then fluency, whether the script and language are what you asked for, and whether names, numbers and technical terms survived.
4. **Compare cost and speed** from `usage` and the run time, alongside quality.
5. **Re-run when you change models or prompts.** Models change often, so keep your prompt set.

Candidates come from [choosing a model](/docs/choosing-a-model) and [/models](/models). Use the exact ids from there, or from `GET /v1/models`, which lists what your key can call.

## A script to compare models on the same prompt

This script sends one Bengali prompt to each model and prints the token usage, the finish reason, the time taken and the answer. Add the model ids you want to compare to `MODELS`, or set them in the `TOKENS_MODELS` environment variable, separated by commas.

```python title="compare_models.py"
import os
import time
import unicodedata

from openai import OpenAI

client = OpenAI(base_url="https://tokens.bd/v1", api_key=os.environ["TOKENS_API_KEY"])

MODELS = os.environ.get("TOKENS_MODELS", "deepseek/deepseek-v4.1-flash").split(",")

SYSTEM = (
    "তুমি একজন সহায়ক সহকারী। সবসময় বাংলা অক্ষরে বাংলায় উত্তর দেবে। "
    "কোড, ভেরিয়েবলের নাম এবং JSON কী ইংরেজিতেই রাখবে।"
)
PROMPT = unicodedata.normalize(
    "NFC",
    "নিচের বার্তাটি দুই বাক্যে সংক্ষেপে লেখো এবং গ্রাহকের জন্য একটি বিনীত উত্তরের খসড়া দাও।\n\n"
    "বার্তা: আমি গতকাল কার্ড দিয়ে টাকা দিয়েছি, কিন্তু আমার অ্যাকাউন্টে ক্রেডিট আসেনি। "
    "পেমেন্টের রসিদ আমার কাছে আছে। দয়া করে দ্রুত দেখুন।",
)

print(f"Prompt: {len(PROMPT)} characters\n")

for model in (m.strip() for m in MODELS if m.strip()):
    started = time.time()
    try:
        resp = client.chat.completions.create(
            model=model,
            messages=[
                {"role": "system", "content": SYSTEM},
                {"role": "user", "content": PROMPT},
            ],
            max_tokens=1000,
        )
    except Exception as err:  # a model your key cannot call, a rate limit, and so on
        print(f"=== {model}\nerror: {err}\n")
        continue

    elapsed = time.time() - started
    choice = resp.choices[0]
    usage = resp.usage
    print(f"=== {model}")
    print(
        f"input {usage.prompt_tokens} tokens, output {usage.completion_tokens} tokens, "
        f"finish_reason={choice.finish_reason}, {elapsed:.1f}s"
    )
    print(choice.message.content)
    print()
```

Run it with your key set:

:::code-tabs

```bash title="macOS / Linux"
export TOKENS_API_KEY="tok_live_your_key"
export TOKENS_MODELS="deepseek/deepseek-v4.1-flash"
python compare_models.py
```

```powershell title="Windows PowerShell"
$env:TOKENS_API_KEY = "tok_live_your_key"
$env:TOKENS_MODELS = "deepseek/deepseek-v4.1-flash"
$env:PYTHONUTF8 = "1"
python compare_models.py
```

:::

On Windows, `PYTHONUTF8=1` makes Python print Bengali correctly in the console; without it you may see an encoding error even though the request worked.

When you read the output, look at `output tokens` next to `finish_reason`. A `length` means the answer was cut off at `max_tokens`, and the models you are comparing may need different limits to finish. The calls are billed like any others; each run of the script above is small, and a cap on the key you test with keeps a mistake cheap. See [API keys](/docs/api-keys).

## Troubleshooting

**The answer is in English although the prompt was Bengali.** Add an explicit instruction to the system prompt, in English and in Bengali, such as `Always answer in Bengali (Bangla script).` Check that no later instruction or example in the conversation is in English.

**The answer is Bengali written in Latin letters.** Ask for Bangla script by name, and show a short example.

**Boxes, `?` or `�` appear in the output.** If it is only on screen, it is a font or console problem. If it is in the saved file or the data, check your stream decoding and any step that cuts bytes.

**The answer stops mid-sentence.** Check `finish_reason`. If it is `length`, raise `max_tokens`. If your wallet is low, Tokens may have lowered it; see [plans, credits and wallet](/docs/plans-and-wallet).

**The bill is higher than the English equivalent.** That is expected from the token counts. Measure it with the first script, choose the model by cost on your own text, and see [cutting your spend](/docs/cutting-costs).

**Comparison or search misses text that looks identical.** Normalize both sides to NFC, and check for look-alike characters such as `।` and `|`.

---
Page: https://tokens.bd/docs/bengali-text
