Skip to content

PHP

Use Tokens from PHP and Laravel: the openai-php/client package, the openai-php/laravel package and plain cURL. Base URL, a first call, streaming, tool calls, timeouts and Tokens error codes.

Works withPHP
On this page

Tokens speaks the OpenAI Chat Completions protocol, so any PHP code that can send an HTTP request can use it. This page covers three ways: the community package openai-php/client in plain PHP, openai-php/laravel inside a Laravel app, and raw cURL with no packages. For Laravel's own agent package, see Laravel AI SDK.

What was checked

Based on the README and source of openai-php/client and openai-php/laravel, both version 0.21.0 (released 17 September 2026, checked October 2026), and on the PHP and Guzzle documentation. The code was checked against the documentation and source, not run end to end against Tokens. These packages are community-maintained, not published by OpenAI or Tokens.

What you need#

  • PHP 8.2 or newer (both packages require php ^8.2) and Composer.
  • A Tokens key from API keys, exported as TOKENS_API_KEY.
  • A model id from /models.

The base URL for every example is https://tokens.bd/v1. It includes /v1. In openai-php/client you pass it to withBaseUri() exactly like that: the library appends a / and then the resource path, so requests go to https://tokens.bd/v1/chat/completions. Do not add a trailing slash or /chat/completions yourself. Write the full URL with https://. If you leave the scheme off, the library adds https:// for you.

openai-php/client#

Install#

bash
composer require openai-php/client guzzlehttp/guzzle
export TOKENS_API_KEY="tok_live_your_key"

The package needs a PSR-18 HTTP client. The README says to allow the php-http/discovery Composer plugin or to install a client such as Guzzle yourself. Installing Guzzle as above is the simplest.

Create the client and make a call#

hello.php
<?php

declare(strict_types=1);

require __DIR__ . '/vendor/autoload.php';

$client = OpenAI::factory()
    ->withApiKey((string) getenv('TOKENS_API_KEY'))
    ->withBaseUri('https://tokens.bd/v1')
    ->withHttpClient(new GuzzleHttp\Client([
        'connect_timeout' => 10,
        'timeout' => 300,
    ]))
    ->make();

$response = $client->chat()->create([
    'model' => 'deepseek/deepseek-v4.1-flash',
    'messages' => [
        ['role' => 'system', 'content' => 'Answer in one short paragraph.'],
        ['role' => 'user', 'content' => 'When should I use a queue instead of running code in the request?'],
    ],
    'max_tokens' => 400,
]);

echo $response->choices[0]->message->content, PHP_EOL;
echo "{$response->usage->promptTokens} in, {$response->usage->completionTokens} out", PHP_EOL;
bash
php hello.php

OpenAI::factory(), withApiKey(), withBaseUri(), withHttpClient() and make() are the factory methods from the README. withHttpHeader() adds a header to every request if you need one. Do not use OpenAI::client($key) for Tokens: it has no base URI argument and would send your key to OpenAI.

Choose the model#

model takes the Tokens id exactly as listed in /models, for example deepseek/deepseek-v4.1-flash. To check ids from code:

php
foreach ($client->models()->list()->data as $model) {
    echo $model->id, PHP_EOL;
}

A wrong id returns 404 model_not_found.

Stream responses#

Use createStreamed() and loop over the result. Ask for usage, or the stream carries no token counts:

php
$stream = $client->chat()->createStreamed([
    'model' => 'deepseek/deepseek-v4.1-flash',
    'messages' => [
        ['role' => 'user', 'content' => 'Write a haiku about merge conflicts.'],
    ],
    'stream_options' => ['include_usage' => true],
]);

foreach ($stream as $chunk) {
    $text = $chunk->choices[0]->delta->content ?? null;
    if ($text !== null) {
        echo $text;
        flush();
    }

    if ($chunk->usage !== null) {
        echo PHP_EOL, "{$chunk->usage->promptTokens} in, {$chunk->usage->completionTokens} out", PHP_EOL;
    }
}

With include_usage on, the last chunk has an empty choices list and carries only usage. The ?? null handles that. usage is null on every other chunk. See Streaming.

If the key or a limit fails during a stream, the library throws OpenAI\Exceptions\ErrorException from inside the foreach, so put the loop inside your try block.

Tool calls#

Tool calling works on models that support it. Check the model's page in /models first.

php
$tools = [[
    'type' => 'function',
    'function' => [
        'name' => 'get_weather',
        'description' => 'Current weather for a city',
        'parameters' => [
            'type' => 'object',
            'properties' => ['city' => ['type' => 'string']],
            'required' => ['city'],
        ],
    ],
]];

$messages = [['role' => 'user', 'content' => 'Is it raining in Dhaka?']];

$first = $client->chat()->create([
    'model' => 'deepseek/deepseek-v4.1-flash',
    'messages' => $messages,
    'tools' => $tools,
]);

$message = $first->choices[0]->message;

if ($message->toolCalls !== []) {
    $messages[] = [
        'role' => 'assistant',
        'content' => $message->content,
        'tool_calls' => array_map(fn ($call) => [
            'id' => $call->id,
            'type' => 'function',
            'function' => ['name' => $call->function->name, 'arguments' => $call->function->arguments],
        ], $message->toolCalls),
    ];

    foreach ($message->toolCalls as $call) {
        $args = json_decode($call->function->arguments, true);
        $result = ['city' => $args['city'], 'condition' => 'light rain', 'temp_c' => 29]; // your real lookup here
        $messages[] = [
            'role' => 'tool',
            'tool_call_id' => $call->id,
            'content' => json_encode($result),
        ];
    }

    $final = $client->chat()->create([
        'model' => 'deepseek/deepseek-v4.1-flash',
        'messages' => $messages,
        'tools' => $tools,
    ]);

    echo $final->choices[0]->message->content, PHP_EOL;
}

The request and response shapes are in Tool calling.

Timeouts for long requests#

The README says the default timeout depends on the HTTP client you use, and the way to raise it is to pass a configured client to withHttpClient(). With Guzzle:

OptionMeaning (from the Guzzle docs)
connect_timeoutSeconds to wait while connecting. Default 0, which waits forever.
timeoutTotal time for the whole request, in seconds. Default 0, which waits forever.
read_timeoutTimeout for individual reads on a streamed body. Defaults to the default_socket_timeout ini setting.

Set timeout above your longest non-streamed answer, as in the first example. For streamed answers, a total timeout also covers the time you spend reading the stream, so a long answer can be cut off. For streams, use a client that has no total limit and relies on read_timeout:

php
$streamClient = OpenAI::factory()
    ->withApiKey((string) getenv('TOKENS_API_KEY'))
    ->withBaseUri('https://tokens.bd/v1')
    ->withHttpClient(new GuzzleHttp\Client([
        'connect_timeout' => 10,
        'timeout' => 0,
        'read_timeout' => 120,
    ]))
    ->make();

On the Tokens side, the gateway waits up to 600 seconds for response headers from the upstream. Streaming means you are not holding a silent connection that long.

PHP can also stop a long request before your client does. The web server setting max_execution_time in php.ini defaults to 30 seconds for web requests (the command line has no limit). For long generations in a web app, stream the answer, or run the call in a queue job.

Handle errors#

The library reads Tokens' OpenAI-style error body, so ErrorException carries the Tokens code:

php
use OpenAI\Exceptions\ErrorException;
use OpenAI\Exceptions\TransporterException;

try {
    $client->chat()->create([
        'model' => 'deepseek/deepseek-v4.1-flash',
        'messages' => [['role' => 'user', 'content' => 'hi']],
    ]);
} catch (ErrorException $e) {
    $requestId = $e->response->getHeaderLine('x-tokens-request-id');
    error_log(sprintf('%d %s %s request=%s', $e->getStatusCode(), $e->getErrorCode(), $e->getErrorMessage(), $requestId));

    if ($e->getErrorCode() === 'insufficient_credits') {
        // top up at https://tokens.bd/dashboard/billing
    }
} catch (TransporterException $e) {
    error_log('Network problem, timeout or server error: ' . $e->getMessage());
}

From the library source, with Guzzle as the HTTP client: a 4xx response with a JSON error body (401, 402, 403, 404 and 429 included) becomes an ErrorException. Network failures, timeouts and 5xx responses become a TransporterException, and the Guzzle exception sits in $e->getPrevious(), which has the response if one arrived. The library also defines RateLimitException and ServerException for HTTP clients that do not throw on error statuses, and each has a public $response.

Log x-tokens-request-id with every failure so support can trace the request. The codes you are likely to meet are invalid_api_key (401), model_not_allowed_on_key and tier_permission_denied (403), insufficient_credits (402), and rate_limited, concurrency_limit and window_exhausted (429). Errors lists them all.

openai-php/laravel#

The Laravel package wraps the same client in a service provider and an OpenAI facade. It requires PHP 8.2+ and Laravel ^11.29, ^12.12 or ^13.0.

Install#

bash
composer require openai-php/laravel
php artisan openai:install

The second command creates config/openai.php and appends blank OPENAI_API_KEY and OPENAI_ORGANIZATION lines to .env. Tokens does not use an organization, so delete the OPENAI_ORGANIZATION line.

Configure#

The README documents these variables:

.env
OPENAI_API_KEY=tok_live_your_key
OPENAI_BASE_URL=https://tokens.bd/v1
OPENAI_REQUEST_TIMEOUT=300
VariableConfig keyNotes
OPENAI_API_KEYapi_keyYour Tokens key.
OPENAI_BASE_URLbase_uriDefaults to api.openai.com/v1. Use https://tokens.bd/v1, with /v1.
OPENAI_REQUEST_TIMEOUTrequest_timeoutSeconds, default 30. Raise it for long answers.

Use Tokens-specific names if anything else reads OPENAI_API_KEY

OPENAI_API_KEY is the variable name many packages read. For example, the built-in openai provider in laravel/ai reads it and sends it to OpenAI by default. If you put a Tokens key there, another package could send it to the wrong place. Safer: edit config/openai.php to read your own names and leave OPENAI_* unset.

config/openai.php
return [
    'api_key' => env('TOKENS_API_KEY'),
    'base_uri' => env('TOKENS_BASE_URL', 'https://tokens.bd/v1'),
    'request_timeout' => env('TOKENS_REQUEST_TIMEOUT', 300),
];

The published file also lists organization and project. Leave them out or leave them null. After changing .env or config on a server that caches config, run php artisan config:clear.

Make a call#

php
use OpenAI\Laravel\Facades\OpenAI;

$response = OpenAI::chat()->create([
    'model' => 'deepseek/deepseek-v4.1-flash',
    'messages' => [
        ['role' => 'user', 'content' => 'Explain job batching in Laravel in two sentences.'],
    ],
]);

echo $response->choices[0]->message->content;

The facade exposes the same chat(), models() and other resources as the client above, so streaming, tool calls and error handling are identical to the previous section. The README's own example uses OpenAI::responses(); Tokens also serves /v1/responses (Responses), but the Chat Completions form above is the one this page checked.

Stream from a route#

routes/web.php
use Illuminate\Support\Facades\Route;
use OpenAI\Laravel\Facades\OpenAI;

Route::get('/ask', function () {
    return response()->stream(function () {
        $stream = OpenAI::chat()->createStreamed([
            'model' => 'deepseek/deepseek-v4.1-flash',
            'messages' => [['role' => 'user', 'content' => 'Explain queues in Laravel.']],
        ]);

        foreach ($stream as $chunk) {
            $text = $chunk->choices[0]->delta->content ?? null;
            if ($text !== null) {
                echo $text;
                if (ob_get_level() > 0) {
                    ob_flush();
                }
                flush();
            }
        }
    }, 200, ['Content-Type' => 'text/plain; charset=utf-8', 'Cache-Control' => 'no-cache']);
});

Timeouts in Laravel#

The package builds its Guzzle client with request_timeout as Guzzle's timeout option. That is a total limit, so it also bounds a streamed answer. Set it above your longest answer. If you need Guzzle's read_timeout instead, build the client with the factory (previous section) and register it in your own service provider in place of the package's.

Raw cURL#

No Composer packages needed. This uses PHP's curl extension.

curl-hello.php
<?php

declare(strict_types=1);

$ch = curl_init('https://tokens.bd/v1/chat/completions');
$requestId = '';

curl_setopt_array($ch, [
    CURLOPT_POST => true,
    CURLOPT_HTTPHEADER => [
        'Authorization: Bearer ' . getenv('TOKENS_API_KEY'),
        'Content-Type: application/json',
    ],
    CURLOPT_POSTFIELDS => json_encode([
        'model' => 'deepseek/deepseek-v4.1-flash',
        'messages' => [['role' => 'user', 'content' => 'Say hello in five words.']],
        'max_tokens' => 100,
    ], JSON_THROW_ON_ERROR),
    CURLOPT_RETURNTRANSFER => true,
    CURLOPT_CONNECTTIMEOUT => 10,
    CURLOPT_TIMEOUT => 300,
    CURLOPT_HEADERFUNCTION => function ($ch, string $header) use (&$requestId): int {
        if (stripos($header, 'x-tokens-request-id:') === 0) {
            $requestId = trim(substr($header, strlen('x-tokens-request-id:')));
        }
        return strlen($header);
    },
]);

$body = curl_exec($ch);

if ($body === false) {
    fwrite(STDERR, 'cURL error ' . curl_errno($ch) . ': ' . curl_error($ch) . PHP_EOL);
    exit(1);
}

$status = curl_getinfo($ch, CURLINFO_RESPONSE_CODE);
curl_close($ch);

$data = json_decode($body, true, flags: JSON_THROW_ON_ERROR);

if ($status >= 400) {
    $error = $data['error'] ?? [];
    fwrite(STDERR, sprintf("%d %s %s request=%s\n", $status, $error['code'] ?? '', $error['message'] ?? '', $requestId));
    exit(1);
}

echo $data['choices'][0]['message']['content'], PHP_EOL;

CURLOPT_TIMEOUT is the total time allowed for the request, and 0 (the default) means no limit. CURLOPT_CONNECTTIMEOUT covers only the connection.

To stream, send "stream": true and read the body as it arrives with CURLOPT_WRITEFUNCTION. The body is a series of data: {...} lines that end with data: [DONE]:

curl-stream.php
<?php

declare(strict_types=1);

$buffer = '';
$errorBody = '';

$ch = curl_init('https://tokens.bd/v1/chat/completions');

curl_setopt_array($ch, [
    CURLOPT_POST => true,
    CURLOPT_HTTPHEADER => [
        'Authorization: Bearer ' . getenv('TOKENS_API_KEY'),
        'Content-Type: application/json',
    ],
    CURLOPT_POSTFIELDS => json_encode([
        'model' => 'deepseek/deepseek-v4.1-flash',
        'messages' => [['role' => 'user', 'content' => 'Write a haiku about merge conflicts.']],
        'stream' => true,
        'stream_options' => ['include_usage' => true],
    ], JSON_THROW_ON_ERROR),
    CURLOPT_CONNECTTIMEOUT => 10,
    CURLOPT_TIMEOUT => 0, // no total limit; streams can be long
    CURLOPT_WRITEFUNCTION => function ($ch, string $data) use (&$buffer, &$errorBody): int {
        if (curl_getinfo($ch, CURLINFO_RESPONSE_CODE) >= 400) {
            $errorBody .= $data; // an error comes back as plain JSON, not as a stream
            return strlen($data);
        }

        $buffer .= $data;
        while (($pos = strpos($buffer, "\n")) !== false) {
            $line = trim(substr($buffer, 0, $pos));
            $buffer = substr($buffer, $pos + 1);

            if (!str_starts_with($line, 'data:')) {
                continue;
            }
            $payload = trim(substr($line, 5));
            if ($payload === '[DONE]') {
                continue;
            }

            $chunk = json_decode($payload, true);
            $text = $chunk['choices'][0]['delta']['content'] ?? '';
            if ($text !== '') {
                echo $text;
                flush();
            }
        }

        return strlen($data);
    },
]);

$ok = curl_exec($ch);
$status = curl_getinfo($ch, CURLINFO_RESPONSE_CODE);

if ($ok === false) {
    fwrite(STDERR, 'cURL error ' . curl_errno($ch) . ': ' . curl_error($ch) . PHP_EOL);
    exit(1);
}
if ($status >= 400) {
    fwrite(STDERR, "HTTP $status: $errorBody" . PHP_EOL);
    exit(1);
}
curl_close($ch);
echo PHP_EOL;

The write function must return the number of bytes it received, or cURL aborts the transfer. The choices[0] lookup uses ?? because the final chunk has no choices when include_usage is on.

Check that it works#

Run any of the programs above. A short answer printed means the key, the URL and the model id are right. To check without PHP:

bash
curl -s https://tokens.bd/v1/models -H "Authorization: Bearer $TOKENS_API_KEY"

Limits#

  • openai-php/client models only the OpenAI API. Anything Tokens adds on top (for example the usage endpoint described in Models and usage) is not a method in the package. Call it with cURL or Guzzle.
  • Tokens limits requests per minute and concurrent requests per account. A queue worker pool that runs many jobs at once can hit 429 concurrency_limit. See Rate limits.
  • Do not call Tokens from browser JavaScript. Tokens sends no CORS headers and the key would be visible. Make the call from PHP and return the result.
  • PHP-FPM and web servers buffer output. Streaming to a browser may need ob_flush(), flush() and, on nginx, response buffering turned off for that route.

Troubleshooting#

SymptomCause and fix
404 unsupported_endpointThe base URI is wrong. Use https://tokens.bd/v1 with /v1 and no /chat/completions on the end.
404 model_not_foundThe model id is wrong. Copy it from /models.
401 missing_api_key or invalid_api_keyThe key was empty or wrong. getenv() returns false if the variable is not set in the PHP process. PHP-FPM often does not see shell variables, so use .env, a config file or clear_env = no. In Laravel, run php artisan config:clear.
To use stream requests you must provide an stream handler closure via the OpenAI factory (exception message)You passed a custom HTTP client that is not Guzzle or Symfony. Use Guzzle, or add withStreamHandler() as the README describes.
cURL error 28: Operation timed out or TransporterException after a long waitYour own timeout is shorter than the response. Raise timeout, or stream.
The page stops partway through a streamThe total timeout or max_execution_time ended the request. Raise both.
402 insufficient_creditsTop up in billing.
429 rate_limited or concurrency_limitWait Retry-After seconds or lower your parallelism. window_exhausted means the plan window is used up; retrying will not help until it resets.

Every Tokens error code is in Errors, with the fix for each in Troubleshooting.

Warning

Keep the key in an environment variable or a secrets manager, never in a committed file. If it leaks, rotate it in /dashboard/keys. The old secret stops working immediately.

Was this page helpful?

Still stuck? Open a support ticket

Need help configuring your agent?

Test your connection with the connection tester, or create an API key.