The Vercel AI SDK talks to Tokens through @ai-sdk/openai-compatible, the provider package for any OpenAI-compatible API. Create the provider once with createOpenAICompatible, and then generateText, streamText and tool calling work the same as they do with the first-party providers.
OpenCode uses this same package under the hood, so if you've set Tokens up there (Tokens CLI does it for you), this will look familiar.
Install the packages#
npm install ai @ai-sdk/openai-compatible zod
export TOKENS_API_KEY="tok_live_your_key"The examples target AI SDK 7 (ai@7, @ai-sdk/openai-compatible@3), which needs Node.js 22 or later. On AI SDK 5 or 6, the code is the same except that you pass the system prompt as system instead of instructions.
Create the Tokens provider with createOpenAICompatible#
import { createOpenAICompatible } from "@ai-sdk/openai-compatible";
export const tokens = createOpenAICompatible({
name: "tokens",
baseURL: "https://tokens.bd/v1",
apiKey: process.env.TOKENS_API_KEY,
includeUsage: true, // ask for token counts on streamed responses
});apiKey is sent as Authorization: Bearer <key>. includeUsage: true makes the provider set stream_options.include_usage on streaming calls. Without it, Tokens streams carry no usage data, and usage ends up empty after a streamText call.
Models are addressed by their Tokens ID: tokens("deepseek/deepseek-v4.1-flash"). Copy IDs from /models or GET /v1/models rather than typing them.
generateText#
import { generateText } from "ai";
import { tokens } from "./lib/tokens";
const { text, usage, finishReason } = await generateText({
model: tokens("deepseek/deepseek-v4.1-flash"),
instructions: "You summarize pull requests for busy reviewers.",
prompt:
"Summarize: refactored auth middleware to use a single session lookup; removed two duplicate DB calls.",
maxOutputTokens: 300,
});
console.log(text);
console.log(finishReason, usage.inputTokens, usage.outputTokens);streamText#
streamText isn't awaited. It returns immediately, and you consume textStream.
import { streamText } from "ai";
import { tokens } from "./lib/tokens";
const result = streamText({
model: tokens("deepseek/deepseek-v4.1-flash"),
prompt: "Explain the difference between TCP and UDP for a junior developer.",
maxOutputTokens: 600,
onError: ({ error }) => console.error(error),
});
for await (const delta of result.textStream) {
process.stdout.write(delta);
}
console.log("\n", await result.usage);Errors that happen during a stream are delivered through onError rather than thrown, so wire it up. Otherwise a 402 insufficient_credits just looks like an empty response.
Stream from a Next.js route handler#
Tokens doesn't allow browser calls (no CORS headers), and you don't want your key in client code anyway. Run streamText in a route handler:
import { streamText } from "ai";
import { tokens } from "@/lib/tokens";
export async function POST(req: Request) {
const { prompt } = (await req.json()) as { prompt?: string };
if (!prompt) return Response.json({ error: "prompt required" }, { status: 400 });
const result = streamText({
model: tokens("deepseek/deepseek-v4.1-flash"),
prompt,
maxOutputTokens: 1024,
abortSignal: req.signal,
});
return result.toTextStreamResponse();
}toTextStreamResponse() sends plain text chunks that you can read with fetch and a stream reader. If your frontend uses useChat from @ai-sdk/react, take messages from the request body instead of prompt, convert them with convertToModelMessages, and return result.toUIMessageStreamResponse(). The Tokens part (the provider) doesn't change.
Choose the model and output limit on the server so visitors can't pick an expensive model for you. Put authentication in front of the route too.
Tool calling#
Tools work on models that support them. Check the model page in /models first.
import { generateText, tool, stepCountIs } from "ai";
import { z } from "zod";
import { tokens } from "./lib/tokens";
const { text, steps } = await generateText({
model: tokens("deepseek/deepseek-v4.1-flash"),
tools: {
getWeather: tool({
description: "Current weather for a city",
inputSchema: z.object({ city: z.string() }),
execute: async ({ city }) => ({ city, condition: "light rain", tempC: 29 }),
}),
},
stopWhen: stepCountIs(3),
prompt: "Do I need an umbrella in Dhaka right now?",
});
console.log(text, `(${steps.length} steps)`);If a model doesn't support tools, the upstream provider rejects the request or the model ignores the tools and answers in text. Switch to a model that lists tool support. Wire-level details are in Tool calling.
Errors#
HTTP errors arrive as APICallError (exported from ai), which carries statusCode, responseHeaders and responseBody. The Tokens error code is inside responseBody as JSON, under error.code.
import { APICallError, generateText } from "ai";
import { tokens } from "./lib/tokens";
try {
await generateText({ model: tokens("deepseek/deepseek-v4.1-flash"), prompt: "hi" });
} catch (err) {
if (APICallError.isInstance(err)) {
console.error(err.statusCode, err.responseHeaders?.["x-tokens-request-id"], err.responseBody);
}
}The AI SDK retries failed calls twice by default (maxRetries). Retrying won't fix a 429 window_exhausted, which only clears when the plan window resets. Codes and fixes are in Errors and Troubleshooting.