Skip to content
Agent Guides10 min read

Running OpenClaw and Hermes Agent on a Budget

Tokens Team
Engineering
3 Oct 2026
On this page

OpenClaw and Hermes Agent are built to run without you watching. That changes the cost picture: a coding session ends when you close the terminal, but a personal agent wired to Telegram or a schedule keeps spending until something stops it. This guide covers running OpenClaw and Hermes Agent on a budget with a pay-per-token gateway: the provider config for each, the limits worth setting before you leave them alone, and which models make sense for which jobs.

/.hermes/config.yaml
model:
  default: deepseek/deepseek-v4.1-flash
  provider: custom
  base_url: https://tokens.bd/v1
  key_env: TOKENS_API_KEY
  context_length: 128000

What OpenClaw and Hermes Agent are#

OpenClaw (formerly Clawdbot, then Moltbot) is an open-source personal assistant and agent gateway. It runs locally as a daemon and talks to you through chat apps and its own UI. Custom model providers are configured in ~/.openclaw/openclaw.json, which is JSON5, so comments and trailing commas are allowed. The gateway watches the file and hot-reloads changes.

Hermes Agent is Nous Research's open-source, self-improving agent for the terminal and messaging. It ships a CLI and TUI plus a messaging gateway for Telegram, Discord and others. Its configuration lives in ~/.hermes/config.yaml, which the docs call "the single source of truth".

Both speak the OpenAI Chat Completions format to custom providers, so both work against https://tokens.bd/v1 with a Tokens key. Both can also use the Anthropic Messages format, but the Chat Completions route is the documented default and the one to start with.

Why always-on agents need a different budget#

The cost of an agent turn is dominated by input, not output. Each turn re-sends the system prompt, the tool definitions and the conversation so far. A reply of fifty words might ride on top of 60,000 input tokens of history. At a list price of $0.30 per million input tokens, that turn costs about $0.018 in input alone; at $2.00 per million, about $0.12. Multiply by every message, scheduled check and retry over a month and the model's input price becomes the number that matters.

Three behaviours make this worse for unattended agents:

  • Loops. A tool that keeps failing, or a model that keeps calling it with the same bad arguments, can burn through turns with nobody there to press Ctrl+C.
  • Growing history. Without compaction, every turn is bigger than the last.
  • Side jobs on the main model. Hermes runs auxiliary tasks (vision, web summarisation, context compression, session titles) on the main model unless you configure them separately. A strong, expensive main model means expensive titles.

None of these are exotic. They're what agents do. So the setup should assume them.

Give each agent its own key and a monthly cap#

On Tokens, a spend cap belongs to an API key. When you create a key in the dashboard you can set a monthly spend cap in USD and an allowed-models list. When the key reaches its cap, requests fail with 403 monthly_spend_cap_exceeded until the next month, and nothing else on your account is affected.

Two details shape how you should use this:

  1. Caps and allow-lists can't be edited after creation. To change them, create a new key and switch the agent over. That's a reason to create keys per agent from the start: one for OpenClaw, one for Hermes, separate from the key your editor uses.
  2. The allow-list is a guard against expensive surprises. Hermes can switch models mid-session with /model. OpenClaw can switch with openclaw models set. If the key only allows two cheap models, a typo or an experiment can't move the agent onto a frontier model.

Plans can also have usage windows (a rolling 5-hour session, weekly, monthly), and you can get dashboard notifications at 50, 75, 90 and 100 percent of a window, plus a low-balance alert. Spend caps for coding agents covers the whole set, including a script that polls your remaining usage.

Configure OpenClaw with a custom OpenAI-compatible provider#

The fastest route is non-interactive onboarding with the official flags:

bash
export CUSTOM_API_KEY="tok_live_your_key"
openclaw onboard --non-interactive --accept-risk --skip-health \
  --mode local \
  --auth-choice custom-api-key \
  --custom-base-url "https://tokens.bd/v1" \
  --custom-model-id "deepseek/deepseek-v4.1-flash" \
  --custom-provider-id "tokens" \
  --custom-compatibility openai

If you leave out --custom-api-key, onboarding reads CUSTOM_API_KEY from the environment, as here. For full control, edit the config file directly:

/.openclaw/openclaw.json
{
  agents: {
    defaults: {
      model: { primary: "tokens/deepseek/deepseek-v4.1-flash" },
    },
  },
  models: {
    mode: "merge",
    providers: {
      tokens: {
        baseUrl: "https://tokens.bd/v1",
        apiKey: "${TOKENS_API_KEY}",
        api: "openai-completions",
        timeoutSeconds: 300,
        models: [
          {
            id: "deepseek/deepseek-v4.1-flash",
            name: "DeepSeek V4.1 Flash (Tokens)",
            input: ["text"],
            contextWindow: 128000, // set from the model's page on /models
            maxTokens: 8192,
          },
        ],
      },
    },
  },
}

The model reference is tokens/deepseek/deepseek-v4.1-flash: provider ID, then the Tokens model ID, slash included. OpenClaw's own docs use the same pattern for LM Studio models with slashes in their IDs.

Two fields deserve attention for cost reasons. If maxTokens is unknown, OpenClaw omits both max_tokens and max_completion_tokens, which means no output cap on any request. Set it. If contextWindow is unknown, OpenClaw assumes 200,000, which may be wrong in either direction and affects when it trims history. Set that too, from the model's page in the catalog.

Check the setup with openclaw models list and openclaw models status. openclaw models status --probe sends a live request, so it costs a few tokens.

Configure Hermes Agent with a custom endpoint#

The docs recommend the wizard. Run it from your shell, not inside a chat session (in a session, /model only switches between providers that already exist):

bash
hermes model
# Choose "Custom endpoint (self-hosted / VLLM / etc.)"
#   API base URL -> https://tokens.bd/v1
#   API key      -> tok_live_your_key
#   Model name   -> deepseek/deepseek-v4.1-flash

Or write the config shown at the top of this post. Put the key in ~/.hermes/.env as TOKENS_API_KEY=tok_live_your_key and reference it with key_env, so the YAML file holds no secret. If you want Tokens as one named provider among several, use the providers: block instead:

/.hermes/config.yaml
providers:
  tokens:
    api: https://tokens.bd/v1
    key_env: TOKENS_API_KEY
    transport: chat_completions
    default_model: deepseek/deepseek-v4.1-flash
    context_length: 128000

Three Hermes requirements to know before you pick a model:

  • At least 64,000 tokens of context. Hermes refuses smaller windows at startup. It prints a "Context limit: X tokens" line when it starts, so you can see what it detected.
  • Set context_length explicitly. Hermes checks your config first, then the endpoint's /models, then its defaults. Tokens' GET /v1/models doesn't include context windows, so the config value is the one that counts.
  • Tool calling is required. The endpoint and model must support OpenAI-style tool calls.

Verify with hermes status, hermes doctor, and a one-shot request: hermes chat --oneshot -q "Reply with OK".

Pick models by job, not by leaderboard#

Personal agents mostly do small things: answer a message, call a tool, summarise a page, decide whether anything needs doing. A frontier model is wasted on most of that. The approach I'd take is a cheap, fast default with a stronger model available for the occasional hard task.

List prices from the makers, per million tokens, checked October 2026:

JobModelContextInput / outputWhy
Default chat and tool callsDeepSeek V4.1 Flash1M$0.30 / $1.20 peak, $0.15 / $0.60 off-peakTools, and a non-thinking mode for quick replies
Default chat and tool callsQwen 3.8 Flash1M$0.15 / $0.47Cheap, fast, tools and image input
Default chat and tool callsMiMo V2.6 Flash1M$0.14 / $0.28Low-cost reasoning model, tools and vision
Short side jobs (titles, short summaries)Qwen 3.7 Flash1M$0.03 / $0.13 up to 32K inputPriced by input size; jumps to $0.10 / $0.40 above 32K
Short side jobsStep 3.5 Flash256K$0.10 / $0.30Text only, tools, fast
Hard multi-step workGLM-5.31M$1.40 / $4.40Coding and agent flagship; thinking always on
Hard multi-step workClaude Sonnet 5.51M$2.00 / $10.00Tools and vision, strong all-rounder

Every model in the table clears Hermes' 64K minimum comfortably. What you pay through Tokens is listed on the model catalog and pricing; the numbers above are the makers' own list prices, useful for comparing models against each other.

A few cautions from the makers' own documentation:

  • Always-on thinking costs output tokens. GLM-5.3 and GLM-5.3 Flash always reason before answering. For an agent that mostly replies "done" or "nothing new", a model with a non-thinking mode is cheaper per turn.
  • GPT-6 Luna looks cheap ($0.10 / $0.50 up to 272K input), but OpenAI documents full tool support in the Responses API, while tools over Chat Completions only work with reasoning_effort set to none. Both agents use Chat Completions by default, so check that before choosing it.
  • Laguna S 2.1 ($0.10 / $0.20) has a known issue with nested JSON array arguments in tool calls, according to Poolside. Agents with complex tool schemas may trip over it.
  • DeepSeek's list price depends on the time of day. Peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays. That's the maker's pricing; check the catalog for what applies on Tokens.

For Hermes, route the auxiliary tasks to the cheapest model you trust under auxiliary.* in the config, even if the main model stays stronger. For both agents, set the output limit per model. For a deeper look at budget models, see cheap coding models that hold up, and to turn per-turn numbers into a monthly figure, estimate your monthly token bill.

Before you leave it running#

A short checklist, in the order I'd do it:

  • A dedicated key per agent, with a monthly spend cap and an allowed-models list
  • maxTokens (OpenClaw) set for every model; context_length (Hermes) set explicitly
  • A cheap default model; the strong model only where you've chosen it
  • Hermes auxiliary tasks pointed at a cheap model
  • Usage alerts and the low-balance alert turned on in Notifications

If you're starting from scratch, create the key at API keys; the OpenClaw and Hermes Agent docs pages have the same configs in reference form.

Third-party details (agent configuration, model specs and list prices) checked on 2026-10-03.

Sources: https://docs.openclaw.ai/concepts/model-providers/custom-providers · https://docs.openclaw.ai/start/wizard-cli-automation · https://docs.openclaw.ai/cli/models · https://hermes-agent.nousresearch.com/docs/integrations/providers · https://hermes-agent.nousresearch.com/docs/user-guide/configuring-models · https://api-docs.deepseek.com/quick_start/pricing · https://www.alibabacloud.com/help/en/model-studio/model-pricing · https://mimo.mi.com · https://platform.stepfun.ai/docs/en/guides/pricing/details · https://docs.z.ai/guides/overview/pricing · https://platform.claude.com/docs/en/about-claude/pricing · https://developers.openai.com/api/docs/pricing · https://poolside.ai/blog/introducing-laguna-s-2-1

Was this page helpful?

Still stuck? Open a support ticket

Use the coding models you already know, through one API

One key for OpenAI- and Anthropic-compatible tools. Pay in BDT or USD, and keep the coding agent you already use.

Create an account