A web page or mobile app cannot call Tokens directly. Anything shipped to a user's device can be read by that user, and a Tokens key spends your money. Browsers add a second block: Tokens does not send CORS headers on its responses, so a page's JavaScript cannot read the answer anyway.
The fix is one small backend of your own. The app calls your backend, your backend checks who is asking and calls Tokens with the key, and the answer goes back to the app. This page shows that backend for Next.js and for Express, with streaming passed through, then the same idea for a mobile app.
Why a key in client code fails#
- The key is not secret. A browser shows every request in its developer tools. A mobile app can be unpacked and its strings read. Environment variables that a framework copies into the client bundle, such as those starting with
NEXT_PUBLIC_in Next.js, are public by definition. - A leaked key is a spending problem. Anyone with it can run requests until the key's spend cap or your balance is gone.
- CORS does not protect you, and does not help you either. The gateway answers the browser's preflight request, but the real responses carry no
Access-Control-Allow-Originheader, so the browser refuses to hand the answer to your page. Do not read that as a safety mechanism: the request can still leave the browser with your key in it. Treat a key in client code as already leaked.
If a key has ever been in client code, revoke it in API keys and create a new one. See API keys.
The pattern#
Browser or app --(your user's session)--> Your backend --(Tokens key)--> Tokens
<------- streamed answer --------------- <---- stream -----Your backend does five things:
- Authenticates your user. Use the session or token your app already has. An open endpoint is an open wallet.
- Validates the input. Accept a message list or a prompt, check its size, and ignore everything else the client sends.
- Fixes the model and limits on the server. The client does not choose
model,max_tokensor tools. Otherwise a user can pick the most expensive model and the largest output. - Calls Tokens with the key from an environment variable, and passes the stream through without buffering.
- Hides upstream errors. A billing or limit error is your problem, not your user's. Log the
x-tokens-request-idand tell the user something short.
Use a dedicated key for this service, with a monthly spend cap and an allowed-models list, so a bug or an abusive user can only cost a known amount. See API keys and the production checklist.
Next.js route handler with streaming#
This is a route handler in the App Router. It runs only on the server, so process.env.TOKENS_API_KEY is never sent to the browser.
import { getSessionUserId } from "@/lib/session"; // your own auth
export const runtime = "nodejs";
export const dynamic = "force-dynamic";
const TOKENS_URL = "https://tokens.bd/v1/chat/completions";
const MODEL = "deepseek/deepseek-v4.1-flash";
type ChatMessage = { role: "user" | "assistant"; content: string };
function parseMessages(value: unknown): ChatMessage[] | null {
if (!Array.isArray(value) || value.length === 0 || value.length > 40) return null;
const messages: ChatMessage[] = [];
for (const item of value) {
if (typeof item !== "object" || item === null) return null;
const { role, content } = item as Record<string, unknown>;
if (role !== "user" && role !== "assistant") return null;
if (typeof content !== "string" || content.length > 20_000) return null;
messages.push({ role, content });
}
return messages;
}
export async function POST(req: Request): Promise<Response> {
const userId = await getSessionUserId(req);
if (!userId) return Response.json({ error: "unauthorized" }, { status: 401 });
const payload: unknown = await req.json().catch(() => null);
const messages = parseMessages((payload as { messages?: unknown } | null)?.messages);
if (!messages) return Response.json({ error: "invalid_request" }, { status: 400 });
let upstream: Response;
try {
upstream = await fetch(TOKENS_URL, {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.TOKENS_API_KEY}`,
"Content-Type": "application/json",
"x-request-id": crypto.randomUUID(),
},
body: JSON.stringify({ model: MODEL, messages, stream: true, max_tokens: 1024 }),
signal: req.signal, // client left: cancel the upstream request too
});
} catch {
return Response.json({ error: "unavailable" }, { status: 502 });
}
if (!upstream.ok || !upstream.body) {
console.error("tokens call failed", {
userId,
status: upstream.status,
requestId: upstream.headers.get("x-tokens-request-id"),
});
await upstream.body?.cancel();
const headers = new Headers();
const retryAfter = upstream.headers.get("retry-after");
if (upstream.status === 429 && retryAfter) headers.set("Retry-After", retryAfter);
return Response.json(
{ error: upstream.status === 429 ? "busy" : "unavailable" },
{ status: upstream.status === 429 ? 429 : 502, headers }
);
}
return new Response(upstream.body, {
headers: {
"Content-Type": "text/event-stream; charset=utf-8",
"Cache-Control": "no-cache, no-transform",
"X-Accel-Buffering": "no",
},
});
}getSessionUserId stands for whatever your app already uses to identify a user (Auth.js, Clerk, Supabase Auth, your own cookie). It is the one line you replace. If your host limits how long a route may run, raise the limit for this route; reasoning models can take minutes.
Read the stream in the browser#
The answer is the same Server-Sent Events text Tokens sends, so the client reads data: lines. Chunks can end in the middle of a line, so keep the unfinished part in a buffer.
export async function streamChat(
messages: { role: "user" | "assistant"; content: string }[],
onText: (text: string) => void,
signal?: AbortSignal
): Promise<void> {
const res = await fetch("/api/chat", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ messages }),
signal,
});
if (!res.ok || !res.body) throw new Error(`Chat failed with status ${res.status}`);
const reader = res.body.pipeThrough(new TextDecoderStream()).getReader();
let buffer = "";
for (;;) {
const { done, value } = await reader.read();
if (done) return;
buffer += value;
const lines = buffer.split("\n");
buffer = lines.pop() ?? "";
for (const line of lines) {
if (!line.startsWith("data:")) continue;
const data = line.slice(5).trim();
if (data === "[DONE]") return;
const delta = JSON.parse(data).choices?.[0]?.delta?.content;
if (delta) onText(delta);
}
}
}Pass an AbortSignal from a "stop" button. The abort cancels your backend request, which cancels the Tokens request, and Tokens bills only what was generated up to that point.
Node.js with Express#
The same backend as a plain Express app. This uses the global fetch of Node.js 18 or later and ES modules.
import express from "express";
import { randomUUID } from "node:crypto";
import { Readable } from "node:stream";
import { pipeline } from "node:stream/promises";
import { requireUser } from "./auth.js"; // your own auth middleware
const TOKENS_URL = "https://tokens.bd/v1/chat/completions";
const MODEL = "deepseek/deepseek-v4.1-flash";
const app = express();
app.use(express.json({ limit: "100kb" }));
function validMessages(messages) {
return (
Array.isArray(messages) &&
messages.length > 0 &&
messages.length <= 40 &&
messages.every(
(m) =>
(m?.role === "user" || m?.role === "assistant") &&
typeof m.content === "string" &&
m.content.length <= 20_000
)
);
}
app.post("/api/chat", requireUser, async (req, res) => {
const messages = req.body?.messages;
if (!validMessages(messages)) return res.status(400).json({ error: "invalid_request" });
// The client closed the connection before the answer finished: stop the upstream request.
const controller = new AbortController();
res.on("close", () => {
if (!res.writableEnded) controller.abort();
});
let upstream;
try {
upstream = await fetch(TOKENS_URL, {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.TOKENS_API_KEY}`,
"Content-Type": "application/json",
"x-request-id": randomUUID(),
},
body: JSON.stringify({ model: MODEL, messages, stream: true, max_tokens: 1024 }),
signal: controller.signal,
});
} catch {
return res.status(502).json({ error: "unavailable" });
}
if (!upstream.ok || !upstream.body) {
console.error("tokens call failed", {
userId: req.user.id,
status: upstream.status,
requestId: upstream.headers.get("x-tokens-request-id"),
});
await upstream.body?.cancel();
const retryAfter = upstream.headers.get("retry-after");
if (upstream.status === 429 && retryAfter) res.set("Retry-After", retryAfter);
return res
.status(upstream.status === 429 ? 429 : 502)
.json({ error: upstream.status === 429 ? "busy" : "unavailable" });
}
res.status(200).set({
"Content-Type": "text/event-stream; charset=utf-8",
"Cache-Control": "no-cache, no-transform",
"X-Accel-Buffering": "no",
});
res.flushHeaders();
try {
await pipeline(Readable.fromWeb(upstream.body), res);
} catch {
// The client left or the upstream stream broke. Nothing more to send.
}
});
app.listen(3001, () => console.log("Listening on http://localhost:3001"));requireUser is your authentication middleware; it sets req.user. The browser code from the previous section works against this server unchanged.
If you run Express behind nginx, set proxy_buffering off; for this route. Without it the stream arrives in one block. Do not compress text/event-stream responses. More in troubleshooting.
Mobile apps#
A mobile app does exactly the same thing: it talks to your backend and never to Tokens. The app sends the user's own session token. Your backend checks it, applies your per-user limits, and calls Tokens with the key.
The simplest mobile design is a non-streaming call. Not every platform's default HTTP client delivers a response body piece by piece, and a plain request is easier to get right first. Add a second backend route that returns the whole answer as JSON:
import { getSessionUserId } from "@/lib/session"; // your own auth
export const runtime = "nodejs";
export const dynamic = "force-dynamic";
export async function POST(req: Request): Promise<Response> {
const userId = await getSessionUserId(req);
if (!userId) return Response.json({ error: "unauthorized" }, { status: 401 });
const payload = (await req.json().catch(() => null)) as { prompt?: unknown } | null;
const prompt = payload?.prompt;
if (typeof prompt !== "string" || prompt.length === 0 || prompt.length > 20_000) {
return Response.json({ error: "invalid_request" }, { status: 400 });
}
let upstream: Response;
try {
upstream = await fetch("https://tokens.bd/v1/chat/completions", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.TOKENS_API_KEY}`,
"Content-Type": "application/json",
"x-request-id": crypto.randomUUID(),
},
body: JSON.stringify({
model: "deepseek/deepseek-v4.1-flash",
messages: [{ role: "user", content: prompt }],
max_tokens: 800,
}),
signal: AbortSignal.timeout(120_000),
});
} catch {
return Response.json({ error: "unavailable" }, { status: 502 });
}
if (!upstream.ok) {
console.error("tokens call failed", {
userId,
status: upstream.status,
requestId: upstream.headers.get("x-tokens-request-id"),
});
return Response.json({ error: "unavailable" }, { status: upstream.status === 429 ? 429 : 502 });
}
const data = (await upstream.json()) as { choices?: { message?: { content?: string } }[] };
return Response.json({ text: data.choices?.[0]?.message?.content ?? "" });
}The app calls that route with its own session token:
import Foundation
struct AskReply: Decodable { let text: String }
func ask(_ prompt: String, sessionToken: String) async throws -> String {
var request = URLRequest(url: URL(string: "https://api.example.com/api/ask")!)
request.httpMethod = "POST"
request.setValue("Bearer \(sessionToken)", forHTTPHeaderField: "Authorization")
request.setValue("application/json", forHTTPHeaderField: "Content-Type")
request.httpBody = try JSONEncoder().encode(["prompt": prompt])
request.timeoutInterval = 130
let (data, response) = try await URLSession.shared.data(for: request)
guard let http = response as? HTTPURLResponse, http.statusCode == 200 else {
throw URLError(.badServerResponse)
}
return try JSONDecoder().decode(AskReply.self, from: data).text
}export async function ask(prompt: string, sessionToken: string): Promise<string> {
const res = await fetch("https://api.example.com/api/ask", {
method: "POST",
headers: {
Authorization: `Bearer ${sessionToken}`,
"Content-Type": "application/json",
},
body: JSON.stringify({ prompt }),
});
if (!res.ok) throw new Error(`Request failed with status ${res.status}`);
const data = (await res.json()) as { text: string };
return data.text;
}Replace https://api.example.com with your own backend's address. The app holds a short-lived session for your user, never a Tokens key, so a stolen phone or a decompiled app costs you one user's session, not your account.
If you want streaming on mobile, first check that your HTTP client exposes the response body as it arrives. If it does not, keep the non-streaming route and show a loading state.
Protect the backend itself#
Moving the key to a server moves the target. Anyone who can call your endpoint can spend through it.
- Rate-limit per user, not per IP only, and cap how many requests run at once. Tokens limits requests per minute and concurrency for your whole account, so one noisy user can use up everyone's share. See rate limits.
- Cap the size of what you accept, as the examples do: number of messages, characters per message and
max_tokens. - Keep the model list on the server. If users can pick a model, choose it from your own list.
- Do not expose upstream error text. Return a short code and keep the details in your logs.
- Use a separate key with a spend cap. If the backend is abused, the cap stops the spending.
- Do not trust the
Originheader as authentication. Any script can set it.
Troubleshooting#
The browser console shows a CORS error. The page is calling Tokens directly. Change it to call your own route, as above.
The answer arrives all at once. Something between your backend and the user buffers the stream: a proxy, a compression layer, or an HTTP client that waits for the full body. Check troubleshooting and test the backend with curl -N.
The backend returns 502 and the log shows 401 or 403. The key is wrong, revoked or restricted. See API keys.
The backend returns 502 and the log shows 402. The balance or plan is used up. Top up in billing. Do not retry in a loop; see errors.
Users see 429 from your backend. Your account hit its per-minute or concurrency limit, or a usage window ran out. The log has the code. See rate limits.