OpenClaw and Hermes Agent are built to run without you watching. That changes the cost picture: a coding session ends when you close the terminal, but a personal agent wired to Telegram or a schedule keeps spending until something stops it. This guide covers running OpenClaw and Hermes Agent on a budget with a pay-per-token gateway: the provider config for each, the limits worth setting before you leave them alone, and which models make sense for which jobs.
model:
default: deepseek/deepseek-v4.1-flash
provider: custom
base_url: https://tokens.bd/v1
key_env: TOKENS_API_KEY
context_length: 128000What OpenClaw and Hermes Agent are#
OpenClaw (formerly Clawdbot, then Moltbot) is an open-source personal assistant and agent gateway. It runs locally as a daemon and talks to you through chat apps and its own UI. Custom model providers are configured in ~/.openclaw/openclaw.json, which is JSON5, so comments and trailing commas are allowed. The gateway watches the file and hot-reloads changes.
Hermes Agent is Nous Research's open-source, self-improving agent for the terminal and messaging. It ships a CLI and TUI plus a messaging gateway for Telegram, Discord and others. Its configuration lives in ~/.hermes/config.yaml, which the docs call "the single source of truth".
Both speak the OpenAI Chat Completions format to custom providers, so both work against https://tokens.bd/v1 with a Tokens key. Both can also use the Anthropic Messages format, but the Chat Completions route is the documented default and the one to start with.
Why always-on agents need a different budget#
The cost of an agent turn is dominated by input, not output. Each turn re-sends the system prompt, the tool definitions and the conversation so far. A reply of fifty words might ride on top of 60,000 input tokens of history. At a list price of $0.30 per million input tokens, that turn costs about $0.018 in input alone; at $2.00 per million, about $0.12. Multiply by every message, scheduled check and retry over a month and the model's input price becomes the number that matters.
Three behaviours make this worse for unattended agents:
- Loops. A tool that keeps failing, or a model that keeps calling it with the same bad arguments, can burn through turns with nobody there to press Ctrl+C.
- Growing history. Without compaction, every turn is bigger than the last.
- Side jobs on the main model. Hermes runs auxiliary tasks (vision, web summarisation, context compression, session titles) on the main model unless you configure them separately. A strong, expensive main model means expensive titles.
None of these are exotic. They're what agents do. So the setup should assume them.
Give each agent its own key and a monthly cap#
On Tokens, a spend cap belongs to an API key. When you create a key in the dashboard you can set a monthly spend cap in USD and an allowed-models list. When the key reaches its cap, requests fail with 403 monthly_spend_cap_exceeded until the next month, and nothing else on your account is affected.
Two details shape how you should use this:
- Caps and allow-lists can't be edited after creation. To change them, create a new key and switch the agent over. That's a reason to create keys per agent from the start: one for OpenClaw, one for Hermes, separate from the key your editor uses.
- The allow-list is a guard against expensive surprises. Hermes can switch models mid-session with
/model. OpenClaw can switch withopenclaw models set. If the key only allows two cheap models, a typo or an experiment can't move the agent onto a frontier model.
Plans can also have usage windows (a rolling 5-hour session, weekly, monthly), and you can get dashboard notifications at 50, 75, 90 and 100 percent of a window, plus a low-balance alert. Spend caps for coding agents covers the whole set, including a script that polls your remaining usage.
Configure OpenClaw with a custom OpenAI-compatible provider#
The fastest route is non-interactive onboarding with the official flags:
export CUSTOM_API_KEY="tok_live_your_key"
openclaw onboard --non-interactive --accept-risk --skip-health \
--mode local \
--auth-choice custom-api-key \
--custom-base-url "https://tokens.bd/v1" \
--custom-model-id "deepseek/deepseek-v4.1-flash" \
--custom-provider-id "tokens" \
--custom-compatibility openaiIf you leave out --custom-api-key, onboarding reads CUSTOM_API_KEY from the environment, as here. For full control, edit the config file directly:
{
agents: {
defaults: {
model: { primary: "tokens/deepseek/deepseek-v4.1-flash" },
},
},
models: {
mode: "merge",
providers: {
tokens: {
baseUrl: "https://tokens.bd/v1",
apiKey: "${TOKENS_API_KEY}",
api: "openai-completions",
timeoutSeconds: 300,
models: [
{
id: "deepseek/deepseek-v4.1-flash",
name: "DeepSeek V4.1 Flash (Tokens)",
input: ["text"],
contextWindow: 128000, // set from the model's page on /models
maxTokens: 8192,
},
],
},
},
},
}The model reference is tokens/deepseek/deepseek-v4.1-flash: provider ID, then the Tokens model ID, slash included. OpenClaw's own docs use the same pattern for LM Studio models with slashes in their IDs.
Two fields deserve attention for cost reasons. If maxTokens is unknown, OpenClaw omits both max_tokens and max_completion_tokens, which means no output cap on any request. Set it. If contextWindow is unknown, OpenClaw assumes 200,000, which may be wrong in either direction and affects when it trims history. Set that too, from the model's page in the catalog.
Check the setup with openclaw models list and openclaw models status. openclaw models status --probe sends a live request, so it costs a few tokens.
Configure Hermes Agent with a custom endpoint#
The docs recommend the wizard. Run it from your shell, not inside a chat session (in a session, /model only switches between providers that already exist):
hermes model
# Choose "Custom endpoint (self-hosted / VLLM / etc.)"
# API base URL -> https://tokens.bd/v1
# API key -> tok_live_your_key
# Model name -> deepseek/deepseek-v4.1-flashOr write the config shown at the top of this post. Put the key in ~/.hermes/.env as TOKENS_API_KEY=tok_live_your_key and reference it with key_env, so the YAML file holds no secret. If you want Tokens as one named provider among several, use the providers: block instead:
providers:
tokens:
api: https://tokens.bd/v1
key_env: TOKENS_API_KEY
transport: chat_completions
default_model: deepseek/deepseek-v4.1-flash
context_length: 128000Three Hermes requirements to know before you pick a model:
- At least 64,000 tokens of context. Hermes refuses smaller windows at startup. It prints a "Context limit: X tokens" line when it starts, so you can see what it detected.
- Set
context_lengthexplicitly. Hermes checks your config first, then the endpoint's/models, then its defaults. Tokens'GET /v1/modelsdoesn't include context windows, so the config value is the one that counts. - Tool calling is required. The endpoint and model must support OpenAI-style tool calls.
Verify with hermes status, hermes doctor, and a one-shot request: hermes chat --oneshot -q "Reply with OK".
Pick models by job, not by leaderboard#
Personal agents mostly do small things: answer a message, call a tool, summarise a page, decide whether anything needs doing. A frontier model is wasted on most of that. The approach I'd take is a cheap, fast default with a stronger model available for the occasional hard task.
List prices from the makers, per million tokens, checked October 2026:
| Job | Model | Context | Input / output | Why |
|---|---|---|---|---|
| Default chat and tool calls | DeepSeek V4.1 Flash | 1M | $0.30 / $1.20 peak, $0.15 / $0.60 off-peak | Tools, and a non-thinking mode for quick replies |
| Default chat and tool calls | Qwen 3.8 Flash | 1M | $0.15 / $0.47 | Cheap, fast, tools and image input |
| Default chat and tool calls | MiMo V2.6 Flash | 1M | $0.14 / $0.28 | Low-cost reasoning model, tools and vision |
| Short side jobs (titles, short summaries) | Qwen 3.7 Flash | 1M | $0.03 / $0.13 up to 32K input | Priced by input size; jumps to $0.10 / $0.40 above 32K |
| Short side jobs | Step 3.5 Flash | 256K | $0.10 / $0.30 | Text only, tools, fast |
| Hard multi-step work | GLM-5.3 | 1M | $1.40 / $4.40 | Coding and agent flagship; thinking always on |
| Hard multi-step work | Claude Sonnet 5.5 | 1M | $2.00 / $10.00 | Tools and vision, strong all-rounder |
Every model in the table clears Hermes' 64K minimum comfortably. What you pay through Tokens is listed on the model catalog and pricing; the numbers above are the makers' own list prices, useful for comparing models against each other.
A few cautions from the makers' own documentation:
- Always-on thinking costs output tokens. GLM-5.3 and GLM-5.3 Flash always reason before answering. For an agent that mostly replies "done" or "nothing new", a model with a non-thinking mode is cheaper per turn.
- GPT-6 Luna looks cheap ($0.10 / $0.50 up to 272K input), but OpenAI documents full tool support in the Responses API, while tools over Chat Completions only work with
reasoning_effortset to none. Both agents use Chat Completions by default, so check that before choosing it. - Laguna S 2.1 ($0.10 / $0.20) has a known issue with nested JSON array arguments in tool calls, according to Poolside. Agents with complex tool schemas may trip over it.
- DeepSeek's list price depends on the time of day. Peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays. That's the maker's pricing; check the catalog for what applies on Tokens.
For Hermes, route the auxiliary tasks to the cheapest model you trust under auxiliary.* in the config, even if the main model stays stronger. For both agents, set the output limit per model. For a deeper look at budget models, see cheap coding models that hold up, and to turn per-turn numbers into a monthly figure, estimate your monthly token bill.
Before you leave it running#
A short checklist, in the order I'd do it:
- A dedicated key per agent, with a monthly spend cap and an allowed-models list
-
maxTokens(OpenClaw) set for every model;context_length(Hermes) set explicitly - A cheap default model; the strong model only where you've chosen it
- Hermes auxiliary tasks pointed at a cheap model
- Usage alerts and the low-balance alert turned on in Notifications
If you're starting from scratch, create the key at API keys; the OpenClaw and Hermes Agent docs pages have the same configs in reference form.
Third-party details (agent configuration, model specs and list prices) checked on 2026-10-03.
Sources: https://docs.openclaw.ai/concepts/model-providers/custom-providers · https://docs.openclaw.ai/start/wizard-cli-automation · https://docs.openclaw.ai/cli/models · https://hermes-agent.nousresearch.com/docs/integrations/providers · https://hermes-agent.nousresearch.com/docs/user-guide/configuring-models · https://api-docs.deepseek.com/quick_start/pricing · https://www.alibabacloud.com/help/en/model-studio/model-pricing · https://mimo.mi.com · https://platform.stepfun.ai/docs/en/guides/pricing/details · https://docs.z.ai/guides/overview/pricing · https://platform.claude.com/docs/en/about-claude/pricing · https://developers.openai.com/api/docs/pricing · https://poolside.ai/blog/introducing-laguna-s-2-1