YouMind
Sign in

How to Turn Hermes Into an IB-Grade Finance Analyst

@gemchange_ltd
ENGLISHJun 04, 2026
867K
184
15
13
646

TL;DR

This technical guide details how to configure the Hermes agent for investment banking-grade analysis, focusing on prompt architecture, deterministic math, and integration with financial data sources like EDGAR and FRED.

Im gonna assume you can already get an LLM to say smart things about a stock. Every mf can. Thats not the job.

The job is the harness, the layer around the model that decides whether your analyst is a desk you trust or a random number generator with good grammar. The model is the most swappable part of the whole thing.

gemchanger - inline image

Three parts:

  1. How to run it properly
  2. How to educate it into a finance machine (the real edge)
  3. The repos and services that extend it.

Read the middle part twice.

Not Financial Advice. Do Your Own Research. My own project - @coldvisionXYZ

PART 1 RUN IT, AND UNDERSTAND WHAT YOU'RE RUNNING

One line. It provisions python, node, git, everything, into ~/.hermes/:

bash
1curl -fsSL https://raw.githubusercontent.com/NousResearch/hermes-agent/main/scripts/install.sh | bash

Linux, macOS, WSL2, Android via Termux (auto-detected), native Windows (early beta, no WSL needed). Python 3.11+. Then hermes setup for the wizard, or hermes model / hermes tools / hermes gateway setup piecemeal. Run it with hermes (classic CLI) or hermes --tui (recommended).

The provider abstraction is the quiet superpower

The same runtime drives chat-completions APIs, Anthropic Messages, Codex Responses, an out-of-process Codex app-server path, and Bedrock.

Tool-call formats and provider quirks get normalized by transport adapters, so at the loop level the model surface looks identical no matter whats behind it.

That means provider failover is real, and swapping models is genuinely zero-code.

Practical consequence for your wallet: you dont run one model.

Under auxiliary in config.yaml, every side task, curator runs, vision, embeddings, title generation, session search, the compression summarizer, can pin its own provider, model, base_url and reasoning effort.

So your expensive reasoning model never burns tokens summarizing old chat or naming a session.

There's also smart_model_routing for sending hard turns to the strong model and everything else to a cheap one.

Cheapest sane setup: Nous Portal (one sub, 300+ models + the Tool Gateway for web/browser/image/TTS) as primary, a cheap model pinned for all auxiliary work, local Ollama/vLLM as a free fallback. If you go local, Ollama defaults to a 4k context and silently truncates, set num_ctx to 64k+ or the agent goes stupid and you'll blame the model.

The thing you must understand before you touch config: the prompt is cache-shaped

This is the invariant the entire system is built to protect, and if you dont get it you'll quietly 10x your own bill.

Hermes composes the system prompt in three tiers:

  1. stable
  2. context
  3. volatile
gemchanger - inline image

Stable carries identity (SOUL.md), tool guidance for enabled tools only, the skills index, environment hints. Context is cwd-derived. Volatile changes turn by turn. The tiering is explicit in code for one reason, prompt-prefix cache validity. Providers cache the prefix of your prompt and bill cached tokens at a fraction. Every time the stable tier changes, you bust that cache and pay full freight on everything.

So Hermes treats altering context as almost sacred. From the agent's own engineering rules: the only time it deliberately alters context is during compression. Slash commands that mutate system-prompt state (installing a skill, toggling a tool) default to deferred invalidation, the change takes effect next session, with an opt-in --now flag when you really need it live. /skills install --now is the canonical "I accept the cache bust" pattern.

Why you care as an analyst: every skill you enable and every tool you turn on lives in that stable tier and costs prefix tokens on every single call. A bloated agent with 60 enabled tools is paying for all 60 descriptions on every turn forever. Lean toolsets and on-demand skills arent tidiness, theyre the cost model.

Compression is lineage, not truncation

When context fills, naive agents drop old turns and get amnesia.

Hermes runs two independent compression layers:

  1. a gateway "session hygiene" safety net that fires around 85% of context on a rough estimate before the agent even runs
  2. the real one, the in-loop ContextCompressor that fires around 50% using accurate API-reported token counts.

It summarizes middle turns into a new message rather than deleting them, and sessions track parent-child lineage for every compression split. The summary becomes part of the transcript; the original is preserved in lineage

gemchanger - inline image

Lossy is fine, lossless would OOM.

Sessions are infrastructure, not just transcripts.

They carry source tags and routing metadata, and there's a model-facing session_search tool plus an FTS5 SQLite index over every past turn. Your agent can recall "what did I conclude about COIN three weeks ago" mid-reasoning, not because you stuffed it in context, but because it can query its own history on demand.

Where it runs

Six terminal backends:

  1. local
  2. Docker
  3. SSH
  4. Singularity
  5. Modal
  6. Daytona

Modal and Daytona are serverless, the environment hibernates idle and wakes on demand, costing almost nothing between runs.

A $5 VPS or a hibernating serverless box is the real "$5". Wire the gateway and the same agent runs across 20+ platforms (Telegram, Discord, Slack, Signal, …) on unified session routing, so you talk to it from your phone while it works on the cloud box.

Sessions survive restarts via SQLite in WAL mode with a custom retry layer for multi-process write contention, so cron jobs and gateway work dont corrupt each other.

PART 2 - EDUCATE IT INTO A FINANCE MACHINE

A plain agent is a brilliant intern with no domain training and no judgment discipline. You literally educate it across four channels: identity, playbooks, memory, and standards. Then it self-educates, and your real job becomes curation.

2a. Identity - SOUL.md is slot #1

SOUL.md is literally the first thing in the system prompt, before tools, before skills, before anything. Its the persona and values layer.

For a finance analyst this is where you set epistemic temperament, not personality fluff.

Things like: "You are a skeptical buy-side analyst. You distrust round numbers and narratives. You never state a figure you did not pull from a tool this session.

You would rather say 'insufficient data' than guess. You treat your own prior conclusions as priors to be updated"

That sits in the cache-stable tier and colors every single turn for free. Most people leave SOUL.md default. For an analyst its your most leveraged 1,500 characters.

2b. Playbooks - skills, and why the description is the real engineering

A skill is a SKILL.md, name + description + procedure, stored in ~/.hermes/skills/.

Only the short descriptions sit in the always-loaded skills index.

The full procedure loads on demand when a task matches. So your skill library can be huge without bloating the prefix.

Which means the description is where the skill succeeds or fails.

Its a router. If the description is vague the agent never loads the skill when it should, or loads it when it shouldnt. Write descriptions like trigger conditions, not summaries.

"Use when the user names a US-listed equity and wants a fundamental read" beats "analyzes stocks."

A real analyst skill, note the standards from 2d baked straight into the procedure:

markdown
1---
2name: equity-snapshot
3description: Use when the user names a US-listed ticker and wants a fundamental + price read with risk flags.
4---
5
6# Equity Snapshot
7
81. Resolve ticker -> CIK via the SEC company_tickers map. NEVER guess a ticker.
92. Pull latest 10-K/10-Q (revenue, debt, FCF) via EDGAR. Record the accession number.
103. NEVER compute ratios yourself. Call the calc tool for P/E, margins, YoY.
114. Cross-check revenue across two sources. If they disagree, REPORT the gap, do not pick one.
125. Output: thesis (2 lines), 3 catalysts, 3 risks, valuation read, confidence 0-1,
13 and "what one data point would flip this." If data is thin, return "pass."

Skills support per-platform enable/disable, conditional activation based on tool availability, and prerequisite validation, so a skill can declare "only available if the EDGAR tool is present."

The format is the open agentskills.io standard, so what you write is portable to other compatible agents.

2c. The self-improvement loop, and the maintenance

The agent writes its own skills.

After solving something through trial and error (typically a task with several tool calls), it saves the working approach as a SKILL.md in agent_created/.

The skill-management tool has six actions and the one to know is patch: a targeted fix, preferred over full edit rewrites because its token-efficient.

The agent literally refines its own playbooks in place during use.

Without maintenance, auto-skills metastasize.

You end up with dozens of narrow, overlapping playbooks burning prefix tokens and polluting the router.

gemchanger - inline image

The Curator handles it, and its mechanics are worth knowing. It is not a cron daemon, it runs on an inactivity check: roughly when 7 days have passed since its last run and the agent's been idle 2+ hours, a background fork spins up with its own prompt cache, never touching the active conversation, and runs an LLM review loop that auto-archives stale skills (to .archive/, restorable, nothing's ever lost).

Config lives under curator in config.yaml: interval_hours, min_idle_hours, stale_after_days, archive_after_days.

Treat auto-skills like intern drafts. hermes curator review to eyeball them, pin the genuinely good ones so they survive, let the rest get archived. Curating this loop is the education.

An un-curated agent gets worse over time, not better.

2d. Standards - the discipline, ranked by how much it saves you

These are the rules you bake into SOUL.md and every SKILL.md. In order of impact.

1. Ban arithmetic. LLMs predict tokens. This is the number one source of confident wrong numbers. The clean implementation uses execute_code (covered in 2e) so a deterministic script does the math and the model only decides inputs and interprets outputs.

A DCF the model "estimated" is fiction; one a script computed and the model stress-tested is product.

2. Receipts on every figure. Every number returns with its source and as-of date attached, never bare. The renderer drops anything without a receipt.

The payoff is adversarial: when two sources disagree, the disagreement is the signal.

3. Point-in-time only. Look-ahead bias is the silent killer of every backtest and every "it called it" screenshot. APIs serve restated, current data by default.

Tag facts with as-of dates and gate the agent to only what was knowable on the decision date.

4. Adversarial by construction. One subagent builds the long case, a second is ordered to kill it, a third reconciles at a stated confidence. The value is entirely in the objectives being genuinely opposed, not in agent count.

Neutral-aligned models matter here because they'll argue the bear case hard instead of hedging it into mush.

5. Licensed to abstain. Force "what one data point would flip this" on every call.

If thats data the agent doesnt have, the answer is no position.

2e. The advanced move: programmatic tool calling collapses pipelines to zero context cost

This is the feature that, once you get it, changes how you build every analyst workflow.

Normally a multi-step pipeline burns an LLM turn per step: "I'll search… now I'll read… now I'll summarize… now I'll write."

Every mechanical step costs inference tokens and pollutes context with intermediate junk.

execute_code kills that

The agent writes one Python script that calls Hermes's own tools over a Unix-domain-socket RPC. The script runs in a child process; tool calls travel over the socket back to the parent and dispatch through the same handler as normal tool calls.

Critically: only the script's print() output returns to the model. Intermediate tool results never enter the context window.

python
1# the agent writes this ONCE; the model is involved only at the decision point
2from hermes_tools import edgar_fetch, market_price, calc
3
4tickers = ["TSLA", "NVDA", "COIN"]
5rows = []
6for t in tickers:
7 f = edgar_fetch(t, form="10-Q") # tool call over RPC
8 px = market_price(t) # tool call over RPC
9 pe = calc("pe", price=px["last"], eps=f["eps"]) # deterministic math, not the LLM
10 rows.append({"ticker": t, "pe": pe, "src": f["accession"], "as_of": f["period_end"]})
11
12print(rows) # ONLY this hits the context window

For a finance analyst this is enormous.

A morning sweep over 20 tickers with three tool calls each is 60 tool calls that, done conventionally, would blow your context and your token budget.

As a script its one turn, one clean printed table, math done deterministically inline. You fold whole pipelines into single inferences and the model only thinks at the points that actually need judgment. Build your recurring analyst workflows as execute_code scripts wrapped in skills. (Linux/macOS only, it needs Unix domain sockets.)

2f. Memory - facts vs. user-model

Hermes's memory is three orthogonal mechanisms (the "3 layers" framing is a teaching simplification, in code theyre independent):

  • MEMORY.md - durable facts. ~2,200 char cap.
  • USER.md - the model of you. ~1,375 char cap.
  • SessionDB - SQLite, WAL, FTS5 over every past turn, queried via session_search.

A MemoryStore reads MEMORY.md and USER.md once at session start and embeds them as a single immutable block in the system prompt.

The agent can write to those files mid-session and the writes hit disk, but the in-prompt copy doesnt change until next session, because changing it would bust the prefix cache (same invariant as everywhere else).

Plus 8 optional external providers (Honcho's dialectic user modeling, Mem0, Hindsight, Supermemory…), single-select, hermes memory setup.

The character caps are a feature. They force you to keep USER.md as high-signal infrastructure, not a dumping ground.

For an analyst, USER.md is your mandate: risk tolerance, horizon, the metrics you actually act on, output format.

MEMORY.md is for durable facts the agent earned ("COIN revenue ties cleaner to the 10-K than to FMP").

SessionDB as your scorekeeper, store every call, score it against outcomes, and the agent updates priors while you learn its calibration, where its sharp and where its chronically too bullish. That calibration map is alpha no retail build bothers to collect.

2g. When prompts arent enough: GAPA, then RL

Before you reach for fine-tuning: Hermes has GAPA, systematic prompt optimization for your SOUL.md, skill instructions, and system prompts, instead of manual trial-and-error tweaking.

Try it when performance plateaus.

Past that, Hermes does batch trajectory generation and exports ShareGPT-format traces for SFT, and there's an RL path (Atropos) to actually fine-tune tool-calling behavior on your own trajectories. This is the real ceiling of "educate it", it ends in training a model on how your analyst works.

PART 3 - EXTEND IT: REPOS, MCP, SERVICES

MCP: separate registration from exposure

Add servers via CLI or config.yaml; the agent lists their tools at startup and registers them alongside built-ins. Key discipline, mirroring the cache logic above: whitelist with tools.include so you only expose what you need, because every exposed tool costs prefix tokens and widens your attack surface.

bash
1hermes mcp add github --command npx --args "-y,@modelcontextprotocol/server-github"
2hermes mcp configure github # toggle individual tools
yaml
1mcp_servers:
2 filesystem:
3 command: npx
4 args: ["-y", "@modelcontextprotocol/server-filesystem", "/home/you/research"]
5 tools: { include: [read_file, list_directory] } # read-only, no writes

/reload-mcp to apply.

Composio MCP is the cheat code, one server, hundreds of SaaS connections; people build finance agents on Hermes pulling market data and writing reports to Google Docs entirely through it. And hermes mcp serve runs Hermes as an MCP server, exposing its session history so Claude Desktop or Cursor can query what your analyst found, the agent becomes a knowledge base your other tools read from.

The data faucets (ranked: spine vs garnish)

Crypto spine:

DefiLlama

(free, no key, no real rate limit, TVL/fees/yields/prices, half the paid dashboards are reskins)

+ Helius

(Solana king, enhanced tx parsing turns byte-soup into readable swaps).

Garnish:

Birdeye (OHLCV+websocket), Jupiter (routing/quotes), Dune+Flipside (SQL on-chain), Bitquery (GraphQL multichain), Nansen/Arkham (labels, paid).

TradFi spine:

FRED

(St Louis Fed, highest-signal free macro, the data the desks actually watch)

+ edgartools

(MIT, no key, typed financials out of 10-Ks with accession numbers).

Garnish:

Finnhub (best free tier), FMP (pre-computed ratios), Polygon/Tiingo (clean historical), yfinance (duct tape, breaks silently), GDELT (global news events).

Trap: free tiers are tight (Alpha Vantage ~2 dozen calls/day). You'll think your code broke. Cache everything, and use credential pools (config.yaml) to rotate across multiple keys automatically and survive rate limits.

Repos worth your time

  • NousResearch/hermes-agent

Read the developer-guide docs: architecture, agent-loop, context-compression-and-caching. The DeepWiki and mudrii/hermes-agent-docs mirrors are good for the internals.

gemchanger - inline image
  • 0xNyk/awesome-hermes-agent

Curated skills, tools, MCP servers, integrations. Start here. Skill hubs: the official Hub (680+ skills, 18 categories), skills.sh, ClawHub.

gemchanger - inline image
  • OpenBB-finance/OpenBB

Open-source Bloomberg that ships MCP servers, so it plugs straight into Hermes and hands the agent a huge market-data + analytics surface through one interface. The highest-leverage single bolt-on.

gemchanger - inline image
  • polakowo/vectorbt

Numba-fast backtesting, thousands of variations in seconds. Wrap in a skill so the agent tests every hypothesis before you trust it.

gemchanger - inline image
  • TauricResearch/TradingAgents (+ auronsun/TradingAgents-crypto)

Multi-agent trading-firm sim. Read it for the structure, then build the leaner bull/bear-subagent version you understand. microsoft/qlib and nautechsystems/nautilus_trader for real systematic strategies. wilsonfreitas/awesome-quant as the library map.

gemchanger - inline image
One-click save

Use YouMind for AI deep reading of viral articles

Save the source, ask focused questions, summarize the argument, and turn a viral article into reusable notes in one AI workspace.

Explore YouMind
For creators

Turn your Markdown into a clean 𝕏 article

When you publish your own long-form writing, images, tables, and code blocks make 𝕏 formatting painful. YouMind turns a full Markdown draft into a clean, ready-to-post 𝕏 article.

Try Markdown to 𝕏

More patterns to decode

Recent viral articles

Explore more viral articles