Tokens speaks the OpenAI Chat Completions protocol, so any PHP code that can send an HTTP request can use it. This page covers three ways: the community package openai-php/client in plain PHP, openai-php/laravel inside a Laravel app, and raw cURL with no packages. For Laravel's own agent package, see Laravel AI SDK.
What was checked
Based on the README and source of openai-php/client and openai-php/laravel, both version 0.21.0 (released 17 September 2026, checked October 2026), and on the PHP and Guzzle documentation. The code was checked against the documentation and source, not run end to end against Tokens. These packages are community-maintained, not published by OpenAI or Tokens.
What you need#
- PHP 8.2 or newer (both packages require
php ^8.2) and Composer. - A Tokens key from API keys, exported as
TOKENS_API_KEY. - A model id from /models.
The base URL for every example is https://tokens.bd/v1. It includes /v1. In openai-php/client you pass it to withBaseUri() exactly like that: the library appends a / and then the resource path, so requests go to https://tokens.bd/v1/chat/completions. Do not add a trailing slash or /chat/completions yourself. Write the full URL with https://. If you leave the scheme off, the library adds https:// for you.
openai-php/client#
Install#
composer require openai-php/client guzzlehttp/guzzle
export TOKENS_API_KEY="tok_live_your_key"The package needs a PSR-18 HTTP client. The README says to allow the php-http/discovery Composer plugin or to install a client such as Guzzle yourself. Installing Guzzle as above is the simplest.
Create the client and make a call#
<?php
declare(strict_types=1);
require __DIR__ . '/vendor/autoload.php';
$client = OpenAI::factory()
->withApiKey((string) getenv('TOKENS_API_KEY'))
->withBaseUri('https://tokens.bd/v1')
->withHttpClient(new GuzzleHttp\Client([
'connect_timeout' => 10,
'timeout' => 300,
]))
->make();
$response = $client->chat()->create([
'model' => 'deepseek/deepseek-v4.1-flash',
'messages' => [
['role' => 'system', 'content' => 'Answer in one short paragraph.'],
['role' => 'user', 'content' => 'When should I use a queue instead of running code in the request?'],
],
'max_tokens' => 400,
]);
echo $response->choices[0]->message->content, PHP_EOL;
echo "{$response->usage->promptTokens} in, {$response->usage->completionTokens} out", PHP_EOL;php hello.phpOpenAI::factory(), withApiKey(), withBaseUri(), withHttpClient() and make() are the factory methods from the README. withHttpHeader() adds a header to every request if you need one. Do not use OpenAI::client($key) for Tokens: it has no base URI argument and would send your key to OpenAI.
Choose the model#
model takes the Tokens id exactly as listed in /models, for example deepseek/deepseek-v4.1-flash. To check ids from code:
foreach ($client->models()->list()->data as $model) {
echo $model->id, PHP_EOL;
}A wrong id returns 404 model_not_found.
Stream responses#
Use createStreamed() and loop over the result. Ask for usage, or the stream carries no token counts:
$stream = $client->chat()->createStreamed([
'model' => 'deepseek/deepseek-v4.1-flash',
'messages' => [
['role' => 'user', 'content' => 'Write a haiku about merge conflicts.'],
],
'stream_options' => ['include_usage' => true],
]);
foreach ($stream as $chunk) {
$text = $chunk->choices[0]->delta->content ?? null;
if ($text !== null) {
echo $text;
flush();
}
if ($chunk->usage !== null) {
echo PHP_EOL, "{$chunk->usage->promptTokens} in, {$chunk->usage->completionTokens} out", PHP_EOL;
}
}With include_usage on, the last chunk has an empty choices list and carries only usage. The ?? null handles that. usage is null on every other chunk. See Streaming.
If the key or a limit fails during a stream, the library throws OpenAI\Exceptions\ErrorException from inside the foreach, so put the loop inside your try block.
Tool calls#
Tool calling works on models that support it. Check the model's page in /models first.
$tools = [[
'type' => 'function',
'function' => [
'name' => 'get_weather',
'description' => 'Current weather for a city',
'parameters' => [
'type' => 'object',
'properties' => ['city' => ['type' => 'string']],
'required' => ['city'],
],
],
]];
$messages = [['role' => 'user', 'content' => 'Is it raining in Dhaka?']];
$first = $client->chat()->create([
'model' => 'deepseek/deepseek-v4.1-flash',
'messages' => $messages,
'tools' => $tools,
]);
$message = $first->choices[0]->message;
if ($message->toolCalls !== []) {
$messages[] = [
'role' => 'assistant',
'content' => $message->content,
'tool_calls' => array_map(fn ($call) => [
'id' => $call->id,
'type' => 'function',
'function' => ['name' => $call->function->name, 'arguments' => $call->function->arguments],
], $message->toolCalls),
];
foreach ($message->toolCalls as $call) {
$args = json_decode($call->function->arguments, true);
$result = ['city' => $args['city'], 'condition' => 'light rain', 'temp_c' => 29]; // your real lookup here
$messages[] = [
'role' => 'tool',
'tool_call_id' => $call->id,
'content' => json_encode($result),
];
}
$final = $client->chat()->create([
'model' => 'deepseek/deepseek-v4.1-flash',
'messages' => $messages,
'tools' => $tools,
]);
echo $final->choices[0]->message->content, PHP_EOL;
}The request and response shapes are in Tool calling.
Timeouts for long requests#
The README says the default timeout depends on the HTTP client you use, and the way to raise it is to pass a configured client to withHttpClient(). With Guzzle:
| Option | Meaning (from the Guzzle docs) |
|---|---|
connect_timeout | Seconds to wait while connecting. Default 0, which waits forever. |
timeout | Total time for the whole request, in seconds. Default 0, which waits forever. |
read_timeout | Timeout for individual reads on a streamed body. Defaults to the default_socket_timeout ini setting. |
Set timeout above your longest non-streamed answer, as in the first example. For streamed answers, a total timeout also covers the time you spend reading the stream, so a long answer can be cut off. For streams, use a client that has no total limit and relies on read_timeout:
$streamClient = OpenAI::factory()
->withApiKey((string) getenv('TOKENS_API_KEY'))
->withBaseUri('https://tokens.bd/v1')
->withHttpClient(new GuzzleHttp\Client([
'connect_timeout' => 10,
'timeout' => 0,
'read_timeout' => 120,
]))
->make();On the Tokens side, the gateway waits up to 600 seconds for response headers from the upstream. Streaming means you are not holding a silent connection that long.
PHP can also stop a long request before your client does. The web server setting max_execution_time in php.ini defaults to 30 seconds for web requests (the command line has no limit). For long generations in a web app, stream the answer, or run the call in a queue job.
Handle errors#
The library reads Tokens' OpenAI-style error body, so ErrorException carries the Tokens code:
use OpenAI\Exceptions\ErrorException;
use OpenAI\Exceptions\TransporterException;
try {
$client->chat()->create([
'model' => 'deepseek/deepseek-v4.1-flash',
'messages' => [['role' => 'user', 'content' => 'hi']],
]);
} catch (ErrorException $e) {
$requestId = $e->response->getHeaderLine('x-tokens-request-id');
error_log(sprintf('%d %s %s request=%s', $e->getStatusCode(), $e->getErrorCode(), $e->getErrorMessage(), $requestId));
if ($e->getErrorCode() === 'insufficient_credits') {
// top up at https://tokens.bd/dashboard/billing
}
} catch (TransporterException $e) {
error_log('Network problem, timeout or server error: ' . $e->getMessage());
}From the library source, with Guzzle as the HTTP client: a 4xx response with a JSON error body (401, 402, 403, 404 and 429 included) becomes an ErrorException. Network failures, timeouts and 5xx responses become a TransporterException, and the Guzzle exception sits in $e->getPrevious(), which has the response if one arrived. The library also defines RateLimitException and ServerException for HTTP clients that do not throw on error statuses, and each has a public $response.
Log x-tokens-request-id with every failure so support can trace the request. The codes you are likely to meet are invalid_api_key (401), model_not_allowed_on_key and tier_permission_denied (403), insufficient_credits (402), and rate_limited, concurrency_limit and window_exhausted (429). Errors lists them all.
openai-php/laravel#
The Laravel package wraps the same client in a service provider and an OpenAI facade. It requires PHP 8.2+ and Laravel ^11.29, ^12.12 or ^13.0.
Install#
composer require openai-php/laravel
php artisan openai:installThe second command creates config/openai.php and appends blank OPENAI_API_KEY and OPENAI_ORGANIZATION lines to .env. Tokens does not use an organization, so delete the OPENAI_ORGANIZATION line.
Configure#
The README documents these variables:
OPENAI_API_KEY=tok_live_your_key
OPENAI_BASE_URL=https://tokens.bd/v1
OPENAI_REQUEST_TIMEOUT=300| Variable | Config key | Notes |
|---|---|---|
OPENAI_API_KEY | api_key | Your Tokens key. |
OPENAI_BASE_URL | base_uri | Defaults to api.openai.com/v1. Use https://tokens.bd/v1, with /v1. |
OPENAI_REQUEST_TIMEOUT | request_timeout | Seconds, default 30. Raise it for long answers. |
Use Tokens-specific names if anything else reads OPENAI_API_KEY
OPENAI_API_KEY is the variable name many packages read. For example, the built-in openai provider in laravel/ai reads it and sends it to OpenAI by default. If you put a Tokens key there, another package could send it to the wrong place. Safer: edit config/openai.php to read your own names and leave OPENAI_* unset.
return [
'api_key' => env('TOKENS_API_KEY'),
'base_uri' => env('TOKENS_BASE_URL', 'https://tokens.bd/v1'),
'request_timeout' => env('TOKENS_REQUEST_TIMEOUT', 300),
];The published file also lists organization and project. Leave them out or leave them null. After changing .env or config on a server that caches config, run php artisan config:clear.
Make a call#
use OpenAI\Laravel\Facades\OpenAI;
$response = OpenAI::chat()->create([
'model' => 'deepseek/deepseek-v4.1-flash',
'messages' => [
['role' => 'user', 'content' => 'Explain job batching in Laravel in two sentences.'],
],
]);
echo $response->choices[0]->message->content;The facade exposes the same chat(), models() and other resources as the client above, so streaming, tool calls and error handling are identical to the previous section. The README's own example uses OpenAI::responses(); Tokens also serves /v1/responses (Responses), but the Chat Completions form above is the one this page checked.
Stream from a route#
use Illuminate\Support\Facades\Route;
use OpenAI\Laravel\Facades\OpenAI;
Route::get('/ask', function () {
return response()->stream(function () {
$stream = OpenAI::chat()->createStreamed([
'model' => 'deepseek/deepseek-v4.1-flash',
'messages' => [['role' => 'user', 'content' => 'Explain queues in Laravel.']],
]);
foreach ($stream as $chunk) {
$text = $chunk->choices[0]->delta->content ?? null;
if ($text !== null) {
echo $text;
if (ob_get_level() > 0) {
ob_flush();
}
flush();
}
}
}, 200, ['Content-Type' => 'text/plain; charset=utf-8', 'Cache-Control' => 'no-cache']);
});Timeouts in Laravel#
The package builds its Guzzle client with request_timeout as Guzzle's timeout option. That is a total limit, so it also bounds a streamed answer. Set it above your longest answer. If you need Guzzle's read_timeout instead, build the client with the factory (previous section) and register it in your own service provider in place of the package's.
Raw cURL#
No Composer packages needed. This uses PHP's curl extension.
<?php
declare(strict_types=1);
$ch = curl_init('https://tokens.bd/v1/chat/completions');
$requestId = '';
curl_setopt_array($ch, [
CURLOPT_POST => true,
CURLOPT_HTTPHEADER => [
'Authorization: Bearer ' . getenv('TOKENS_API_KEY'),
'Content-Type: application/json',
],
CURLOPT_POSTFIELDS => json_encode([
'model' => 'deepseek/deepseek-v4.1-flash',
'messages' => [['role' => 'user', 'content' => 'Say hello in five words.']],
'max_tokens' => 100,
], JSON_THROW_ON_ERROR),
CURLOPT_RETURNTRANSFER => true,
CURLOPT_CONNECTTIMEOUT => 10,
CURLOPT_TIMEOUT => 300,
CURLOPT_HEADERFUNCTION => function ($ch, string $header) use (&$requestId): int {
if (stripos($header, 'x-tokens-request-id:') === 0) {
$requestId = trim(substr($header, strlen('x-tokens-request-id:')));
}
return strlen($header);
},
]);
$body = curl_exec($ch);
if ($body === false) {
fwrite(STDERR, 'cURL error ' . curl_errno($ch) . ': ' . curl_error($ch) . PHP_EOL);
exit(1);
}
$status = curl_getinfo($ch, CURLINFO_RESPONSE_CODE);
curl_close($ch);
$data = json_decode($body, true, flags: JSON_THROW_ON_ERROR);
if ($status >= 400) {
$error = $data['error'] ?? [];
fwrite(STDERR, sprintf("%d %s %s request=%s\n", $status, $error['code'] ?? '', $error['message'] ?? '', $requestId));
exit(1);
}
echo $data['choices'][0]['message']['content'], PHP_EOL;CURLOPT_TIMEOUT is the total time allowed for the request, and 0 (the default) means no limit. CURLOPT_CONNECTTIMEOUT covers only the connection.
To stream, send "stream": true and read the body as it arrives with CURLOPT_WRITEFUNCTION. The body is a series of data: {...} lines that end with data: [DONE]:
<?php
declare(strict_types=1);
$buffer = '';
$errorBody = '';
$ch = curl_init('https://tokens.bd/v1/chat/completions');
curl_setopt_array($ch, [
CURLOPT_POST => true,
CURLOPT_HTTPHEADER => [
'Authorization: Bearer ' . getenv('TOKENS_API_KEY'),
'Content-Type: application/json',
],
CURLOPT_POSTFIELDS => json_encode([
'model' => 'deepseek/deepseek-v4.1-flash',
'messages' => [['role' => 'user', 'content' => 'Write a haiku about merge conflicts.']],
'stream' => true,
'stream_options' => ['include_usage' => true],
], JSON_THROW_ON_ERROR),
CURLOPT_CONNECTTIMEOUT => 10,
CURLOPT_TIMEOUT => 0, // no total limit; streams can be long
CURLOPT_WRITEFUNCTION => function ($ch, string $data) use (&$buffer, &$errorBody): int {
if (curl_getinfo($ch, CURLINFO_RESPONSE_CODE) >= 400) {
$errorBody .= $data; // an error comes back as plain JSON, not as a stream
return strlen($data);
}
$buffer .= $data;
while (($pos = strpos($buffer, "\n")) !== false) {
$line = trim(substr($buffer, 0, $pos));
$buffer = substr($buffer, $pos + 1);
if (!str_starts_with($line, 'data:')) {
continue;
}
$payload = trim(substr($line, 5));
if ($payload === '[DONE]') {
continue;
}
$chunk = json_decode($payload, true);
$text = $chunk['choices'][0]['delta']['content'] ?? '';
if ($text !== '') {
echo $text;
flush();
}
}
return strlen($data);
},
]);
$ok = curl_exec($ch);
$status = curl_getinfo($ch, CURLINFO_RESPONSE_CODE);
if ($ok === false) {
fwrite(STDERR, 'cURL error ' . curl_errno($ch) . ': ' . curl_error($ch) . PHP_EOL);
exit(1);
}
if ($status >= 400) {
fwrite(STDERR, "HTTP $status: $errorBody" . PHP_EOL);
exit(1);
}
curl_close($ch);
echo PHP_EOL;The write function must return the number of bytes it received, or cURL aborts the transfer. The choices[0] lookup uses ?? because the final chunk has no choices when include_usage is on.
Check that it works#
Run any of the programs above. A short answer printed means the key, the URL and the model id are right. To check without PHP:
curl -s https://tokens.bd/v1/models -H "Authorization: Bearer $TOKENS_API_KEY"Limits#
openai-php/clientmodels only the OpenAI API. Anything Tokens adds on top (for example the usage endpoint described in Models and usage) is not a method in the package. Call it with cURL or Guzzle.- Tokens limits requests per minute and concurrent requests per account. A queue worker pool that runs many jobs at once can hit
429 concurrency_limit. See Rate limits. - Do not call Tokens from browser JavaScript. Tokens sends no CORS headers and the key would be visible. Make the call from PHP and return the result.
- PHP-FPM and web servers buffer output. Streaming to a browser may need
ob_flush(),flush()and, on nginx, response buffering turned off for that route.
Troubleshooting#
| Symptom | Cause and fix |
|---|---|
404 unsupported_endpoint | The base URI is wrong. Use https://tokens.bd/v1 with /v1 and no /chat/completions on the end. |
404 model_not_found | The model id is wrong. Copy it from /models. |
401 missing_api_key or invalid_api_key | The key was empty or wrong. getenv() returns false if the variable is not set in the PHP process. PHP-FPM often does not see shell variables, so use .env, a config file or clear_env = no. In Laravel, run php artisan config:clear. |
To use stream requests you must provide an stream handler closure via the OpenAI factory (exception message) | You passed a custom HTTP client that is not Guzzle or Symfony. Use Guzzle, or add withStreamHandler() as the README describes. |
cURL error 28: Operation timed out or TransporterException after a long wait | Your own timeout is shorter than the response. Raise timeout, or stream. |
| The page stops partway through a stream | The total timeout or max_execution_time ended the request. Raise both. |
402 insufficient_credits | Top up in billing. |
429 rate_limited or concurrency_limit | Wait Retry-After seconds or lower your parallelism. window_exhausted means the plan window is used up; retrying will not help until it resets. |
Every Tokens error code is in Errors, with the fix for each in Troubleshooting.
Warning
Keep the key in an environment variable or a secrets manager, never in a committed file. If it leaks, rotate it in /dashboard/keys. The old secret stops working immediately.