Skip to content

Node.js and TypeScript (OpenAI SDK)

Use the official openai npm package with Tokens from Node.js and TypeScript: client setup, streaming, error handling with APIError, and a Next.js route handler that keeps your key on the server.

Works withOpenAI Node SDKNext.js
On this page

The official openai npm package talks to Tokens with two settings changed: baseURL and apiKey. This page covers a TypeScript setup, streaming, reading Tokens error codes from APIError, and a Next.js route handler so your key never reaches the browser.

Install the OpenAI Node.js SDK#

bash
npm install openai
export TOKENS_API_KEY="tok_live_your_key"

The examples use ES modules and top-level await, so run them as .mjs or .ts files (for example with npx tsx hello.ts). They assume a current major version of openai (v5 or later).

lib/tokens.ts
import OpenAI from "openai";

export const tokens = new OpenAI({
  baseURL: "https://tokens.bd/v1",
  apiKey: process.env.TOKENS_API_KEY,
  timeout: 120_000, // ms; SDK default is 10 minutes
  maxRetries: 2, // SDK default; retries connection errors, 408, 409, 429, 5xx
});
hello.ts
import { tokens } from "./lib/tokens";

const completion = await tokens.chat.completions.create({
  model: "deepseek/deepseek-v4.1-flash",
  messages: [
    { role: "system", content: "Reply in one short paragraph." },
    { role: "user", content: "What's the difference between interface and type in TypeScript?" },
  ],
  max_tokens: 400,
});

console.log(completion.choices[0]?.message.content);
console.log(completion.usage);

If you'd rather not touch code, the SDK also reads OPENAI_BASE_URL and OPENAI_API_KEY from the environment. Keep in mind those variables affect every OpenAI-based tool in that environment.

Copy model IDs from /models or await tokens.models.list(). Don't guess them.

Stream chat completions#

ts
const stream = await tokens.chat.completions.create({
  model: "deepseek/deepseek-v4.1-flash",
  messages: [{ role: "user", content: "List five git commands I should know, one per line." }],
  stream: true,
  stream_options: { include_usage: true },
});

for await (const chunk of stream) {
  const text = chunk.choices[0]?.delta?.content;
  if (text) process.stdout.write(text);
  if (chunk.usage) {
    console.log(`\n${chunk.usage.prompt_tokens} in, ${chunk.usage.completion_tokens} out`);
  }
}

The last chunk has an empty choices array when include_usage is on, which is why the optional chaining is there. To stop early, break out of the loop or call stream.controller.abort(). More in Streaming.

Handle errors with APIError#

Non-2xx responses throw a subclass of OpenAI.APIError. Tokens puts a specific code on every error, and that tells you far more than the HTTP status.

ts
import OpenAI from "openai";
import { tokens } from "./lib/tokens";

try {
  await tokens.chat.completions.create({
    model: "deepseek/deepseek-v4.1-flash",
    messages: [{ role: "user", content: "hi" }],
  });
} catch (err) {
  if (err instanceof OpenAI.APIConnectionError) {
    console.error("Network problem:", err.message);
  } else if (err instanceof OpenAI.APIError) {
    const requestId = err.headers?.get("x-tokens-request-id");
    console.error(err.status, err.code, err.message, requestId);

    switch (err.code) {
      case "insufficient_credits":
        // 402: top up at /dashboard/billing
        break;
      case "window_exhausted":
        // 429: plan window used up; Retry-After header = seconds until reset
        console.error("Resets in", err.headers?.get("retry-after"), "s");
        break;
      case "model_not_allowed_on_key":
      case "tier_permission_denied":
        // 403: key allow-list or plan doesn't cover this model
        break;
    }
  } else {
    throw err;
  }
}

Two things to know:

  • err.headers is a standard Headers object in current SDK versions, so use .get(). Always log x-tokens-request-id. Support needs it to find your request.
  • The SDK automatically retries 429 and 5xx. The gateway has already tried failing over to another upstream source before it returns 502, 503 or 504. Retrying a window_exhausted won't help until Retry-After has passed, and that can be hours, so set maxRetries low if you'd rather fail fast.

Every code is listed in Errors, and Troubleshooting gives the fix for each.

Keep the key on the server: a Next.js route handler#

Tokens doesn't send CORS headers, so calling it straight from browser JavaScript fails, and it would expose your key anyway. Put the call in a route handler and have your frontend call that.

app/api/chat/route.ts
import OpenAI from "openai";

export const runtime = "nodejs";

const tokens = new OpenAI({
  baseURL: "https://tokens.bd/v1",
  apiKey: process.env.TOKENS_API_KEY, // server-only: no NEXT_PUBLIC_ prefix
});

type ChatMessage = { role: "user" | "assistant"; content: string };

export async function POST(req: Request) {
  const body = (await req.json()) as { messages?: ChatMessage[] };
  if (!Array.isArray(body.messages) || body.messages.length === 0) {
    return Response.json({ error: "messages required" }, { status: 400 });
  }

  try {
    const stream = await tokens.chat.completions.create(
      {
        model: "deepseek/deepseek-v4.1-flash", // pick on the server, not from the client
        messages: body.messages.slice(-20),
        max_tokens: 1024,
        stream: true,
      },
      { signal: req.signal } // stop paying for tokens if the user navigates away
    );

    const encoder = new TextEncoder();
    const text = new ReadableStream<Uint8Array>({
      async start(controller) {
        try {
          for await (const chunk of stream) {
            const delta = chunk.choices[0]?.delta?.content;
            if (delta) controller.enqueue(encoder.encode(delta));
          }
          controller.close();
        } catch (e) {
          controller.error(e);
        }
      },
    });

    return new Response(text, {
      headers: { "Content-Type": "text/plain; charset=utf-8", "Cache-Control": "no-store" },
    });
  } catch (err) {
    if (err instanceof OpenAI.APIError) {
      return Response.json(
        { error: err.code ?? "upstream_error", requestId: err.headers?.get("x-tokens-request-id") },
        { status: err.status ?? 502 }
      );
    }
    throw err;
  }
}

The client side is a plain fetch that reads the body as a stream:

ts
const res = await fetch("/api/chat", {
  method: "POST",
  headers: { "Content-Type": "application/json" },
  body: JSON.stringify({ messages: [{ role: "user", content: "Hello" }] }),
});
const reader = res.body!.pipeThrough(new TextDecoderStream()).getReader();
for (;;) {
  const { value, done } = await reader.read();
  if (done) break;
  console.log(value);
}

A few choices in that handler are deliberate. The model and max_tokens are set on the server so a visitor can't switch your app to an expensive model. Message history is trimmed. You should also add your own authentication and per-user rate limiting in front of this route, because anyone who can reach it is spending your balance. A key with a monthly spend cap (API keys) puts a hard ceiling on the damage.

If you're on the Vercel AI SDK, Vercel AI SDK shows the same pattern with streamText.

Was this page helpful?

Still stuck? Open a support ticket

Need help configuring your agent?

Test your connection with the connection tester, or create an API key.