curl https://app.exnai.com/v1/modelsfrom openai import OpenAI
client = OpenAI(base_url="https://app.exnai.com/v1")Tip
Use our regional edge endpoints to achieve sub-50ms TTFT.
ff
curl https://app.exnai.com/v1/modelsfrom openai import OpenAI
client = OpenAI(base_url="https://app.exnai.com/v1")Tip
Use our regional edge endpoints to achieve sub-50ms TTFT.
ff
Was this page helpful?
Still stuck? Open a support ticket
One key for OpenAI- and Anthropic-compatible tools. Pay in BDT or USD, and keep the coding agent you already use.
Create an accountMillion-token context lets an agent hold a whole repository and a long session. It doesn't make each turn cheaper or faster, and some providers raise the price above 200K, 272K or 512K tokens. Here's the arithmetic.
How prompt caching is priced at DeepSeek, Anthropic, OpenAI, Moonshot, Alibaba and others, why coding agents benefit more than chat apps, what quietly breaks cache hits, and what you can and can't control through a gateway.
Three request formats, one job. Where the system prompt goes, how tool calls and their results are shaped, what the streaming events look like, why output limits and usage fields differ, how a gateway serves the same model in more than one format, and which coding agents speak which.