Hermes Agent is Nous Research's open-source, self-improving agent with a CLI, a TUI and a messaging gateway for Telegram, Discord and others. It connects to Tokens as a custom endpoint over OpenAI-style chat completions at https://tokens.bd/v1.
Hermes needs at least 64K tokens of context
Hermes refuses to start with a model whose context window is under 64,000 tokens, and it needs a model with OpenAI-style tool calling. Check the context window on the model's page in /models before you pick one.
The Tokens CLI doesn't configure Hermes, so use Hermes's own wizard or edit its config file.
Install Hermes Agent#
curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash
source ~/.bashrc # or ~/.zshrciex (irm https://hermes-agent.nousresearch.com/install.ps1)Create a key at /dashboard/keys if you don't have one. API keys explains spend caps and model allow-lists.
Option A: the hermes model wizard#
The Hermes docs recommend the interactive wizard. Run it from your terminal, outside a chat session:
hermes modelChoose Custom endpoint (self-hosted / VLLM / etc.) and enter:
| Prompt | Value |
|---|---|
| API base URL | https://tokens.bd/v1 |
| API key | your tok_live_... key |
| Model name | deepseek/deepseek-v4.1-flash |
Afterwards, open ~/.hermes/config.yaml and add context_length as shown below, so Hermes doesn't have to guess.
Option B: edit ~/.hermes/config.yaml#
The Hermes docs call config.yaml the single source of truth. Put the key in ~/.hermes/.env and reference it by name:
TOKENS_API_KEY=tok_live_your_keyIf you'd rather export it in your shell, that works too:
export TOKENS_API_KEY="tok_live_your_key"$env:TOKENS_API_KEY = "tok_live_your_key" # this window
setx TOKENS_API_KEY "tok_live_your_key" # new windowsThen set the model block:
model:
default: deepseek/deepseek-v4.1-flash
provider: custom
base_url: https://tokens.bd/v1
key_env: TOKENS_API_KEY
context_length: 128000 # placeholder: use the model's real window, at least 64000context_length is a placeholder. Copy the real value from /models. Hermes looks for the context size in your config first, then the endpoint's model list, then its own defaults, so setting it explicitly removes the guesswork.
Keep Tokens as one named provider#
If you use Hermes with several providers, declare Tokens under providers: instead. The older custom_providers: list migrates to this form automatically.
providers:
tokens:
api: https://tokens.bd/v1
key_env: TOKENS_API_KEY
transport: chat_completions
default_model: deepseek/deepseek-v4.1-flash
context_length: 128000 # placeholderHermes also has an anthropic_messages transport. We haven't confirmed whether it appends /v1/messages to the base URL itself, so stick with chat_completions unless you have a reason to change it.
Switch models#
Change model.default (or default_model under your named provider) to another id from /models or GET /v1/models, and update context_length to match the new model. Run hermes model again if you prefer the wizard.
Inside a session, the documented syntax for a named provider is:
/model custom:tokens:deepseek/deepseek-v4.1-flashThe format is /model custom:<name>:<model>. Hermes's docs don't say how it parses a model id that itself contains a slash, so if this doesn't switch, change the model in config.yaml instead. /model inside a session can only switch between providers that are already configured. Choosing a model helps you pick one.
Verify it works#
hermes status
hermes doctor
hermes chat --oneshot -q "Reply with OK"The last command answers once and exits. When Hermes starts a session it prints a Context limit: X tokens line; check that it matches the model's real window.
Troubleshooting#
Hermes refuses to start because of context size. The configured or detected context is under 64,000 tokens. Pick a model with a larger window, and set context_length explicitly.
OPENAI_BASE_URL or LLM_MODEL has no effect. LLM_MODEL was removed. OPENAI_BASE_URL is only read for the openai-api provider, not custom. CUSTOM_BASE_URL survives only as a legacy fallback. Set everything in config.yaml.
401 missing_api_key. key_env names a variable Hermes can't see. Check the spelling in ~/.hermes/.env, or that it's exported in the shell that started Hermes.
Tool calls fail or the agent loops. Hermes depends on OpenAI tool calling. Switch to another model; Choosing a model covers what to look for in an agent model.
Unexpected usage from side tasks. Vision, web summarization, compression and title generation use the main model by default. You can point them at a different model under auxiliary.* in config.yaml.
403 model_not_allowed_on_key or 429 window_exhausted. These are key and plan limits. See Troubleshooting for every error code.