Skip to content
Bot jobsJob breakdowns

Which agent stack when: Hermes, Grok Bot, OpenClaw, Omarchy

People keep asking me which agent stack to pick. Hermes, Grok Bot, OpenClaw, or Omarchy. Like they are four SKUs of the same product. They are not. One is an operating system. Two are self-hosted

Mike GannottiImported from X6 min read
MichaelGannottix article
See this runHouse 313 · 00444

Article

Job breakdowns

People keep asking me which agent stack to pick. Hermes, Grok Bot, OpenClaw, or Omarchy. Like they are four SKUs of the same product. They are not. One is an operating system. Two are self-hosted agent runtimes with very different incentives. One is a managed teammate layer with a computer you do not have to keep awake.

I run all four in some form. Not because I like collecting logos. Because they occupy different jobs. If you pick the wrong layer, you spend a month fighting the product instead of shipping.

Start with the one that is not an agent at all. Omarchy is DHH's Arch Linux desktop, rebuilt around agents as first-class citizens. Quattro, the 4.0 line, put the whole shell in Quickshell. On first boot it asks you to pick a default agent. Super plus Shift plus Ctrl plus A launches it. A crash toast can hand the core dump to that agent. The Install menu will put Hermes Desktop, OpenClaw, and Grok Bot on the same machine if you want them there. The OS is the desk. The other three sit on it.

I pick Omarchy when I want the computer itself to be agent-malleable. Keyboard-first. Files and configs an agent can actually read and edit. I do not pick it because I need a chat app. I pick it because I want Linux that treats Claude Code, Codex, Hermes, Grok CLI, and the rest as system tools, not browser tabs. If you are not ready to live on a Hyprland desktop, do not force it. Omarchy M is the Apple Silicon path, with first-release targeting on M1 and M2. Vintage Intel Macs still have a crew. Windows users can try it in a VM before they partition anything.

Here is the map I actually use. It is conceptual, not a scoreboard. Left is local-first and self-hosted. Right is managed cloud. Bottom is the operating environment. Top is an autonomous agent runtime that keeps going when you close the laptop.

Hermes, from Nous Research, is the self-improving runtime. MIT licensed. Provider-agnostic. You can point it at Nous Portal, OpenRouter, OpenAI, a local endpoint, and switch with hermes model. They pitch a built-in learning loop: it creates skills from experience, improves them in use, searches its own past conversations, and builds a model of you across sessions. It lives on Telegram, Discord, Slack, WhatsApp, Signal, and a real TUI. Terminal work can run local, in Docker, over SSH, in Singularity, or on Modal, Daytona, and Vercel Sandbox, so the agent is not glued to your laptop. Bot Mode in the desktop app makes named agents with group chats. That is the multi-agent shape a lot of people were waiting for.

I pick Hermes when I want an agent that gets sharper on my machine, with any model I am willing to pay for, and a learning loop I can inspect. Homelab. A cheap VPS. A serverless backend that sleeps when idle. I pick it when I care about skills that persist, cron that remembers yesterday, and the ability to steer a subagent mid-flight. I do not pick it when I refuse to operate software. It is a serious runtime. You will touch config.yaml

OpenClaw is the other self-hosted lane, and it is not a Hermes clone. It is a gateway. One always-on process owns channels, sessions, tools, and policy. You talk to it from WhatsApp, Telegram, Discord, iMessage, Slack, Signal, and a long list of others. Desktop apps exist for Mac, Windows, and Linux. The architecture pitch is the one I respect: a trusted Gateway, untrusted execution you can move into a sandbox, a node, or a throwaway cloud machine, and policy enforced in code instead of hoped for in a system prompt. It can embed vendor harnesses as plugins, so Codex, Copilot, or Claude Code run their own loops while OpenClaw keeps channels and state. Stewardship is an independent 501(c)(3). MIT. No paid tier. No hosted product. Donors fund the Foundation. They do not own the roadmap.

I pick OpenClaw when I want the control plane on hardware I own, on channels I already live in, with a governance story I can take to a security review. I pick it for a shared team gateway, live sessions other people can steer, and the option to harden later with sandbox mode, roles, and secret refs. I pick it when independent stewardship matters more than a company's Portal. I do not pick it if I want zero ops. Sandboxing is off by default. Default OpenClaw is a trusted single-operator assistant. Hardening is a choice you make on purpose.

Hermes versus OpenClaw is the comparison people get stuck on. Both self-host. Both do channels, skills, memory, cron. The split is incentive and architecture. Hermes is a Python agent with a learning loop from a venture-backed lab, optional Nous Portal, seven execution backends, and Bot Mode as a product surface. OpenClaw is a Node gateway from a foundation, policy-as-code, vendor harnesses as plugins, and a team multiplayer model. If you want the agent to grow its own skills and swap models without ceremony, I lean Hermes. If you want a Switzerland control plane, signed releases, and channels as the product, I lean OpenClaw. Plenty of serious operators run both. I have.

Grok Bot is the managed answer to the same class of work. SpaceXAI built it as always-on teammates with a persistent cloud computer: browser, filesystem, terminal, plugins, MCP, computer use for sites with no clean API. You message it like a colleague. It keeps going when the laptop closes. Multiple Bots share one user-scoped computer, message each other, sit in group chats, and only pull you in for judgment. Skills and routines turn a demonstrated workflow into something that runs on a schedule. For engineering, the pattern I trust is outer loop and inner loop. Grok Bot gathers context from Slack, GitHub, and docs, then hands a clean prompt to a Cursor cloud agent that actually builds. Grok Bot usage is separate from your Grok and Cursor plan usage. The cloud agents it spawns still draw from your Cursor allowance.

I pick Grok Bot when I do not want to be the SRE for my own agent. When work has to survive a closed laptop, a phone kickoff, and an approval from the grocery line. When I want specialist Bots that can talk to each other without me relaying. When I need the Cursor integration, Auto Review, and, for orgs, Firecracker isolation plus network and audit controls. I do not pick it when I need the runtime on my metal, when I refuse a subscription, or when I need a security boundary between Bots. Docs are explicit: one computer per user, shared files, shared logins, shared sessions. Treat a credential on that machine as visible to every Bot you run.

How I actually decide is boring, which is why it works.

If the job is the computer itself, Omarchy. If the job is a self-hosted agent that learns and is not glued to one model vendor, Hermes. If the job is a self-hosted gateway on my channels, with independent stewardship and a trust boundary I can configure, OpenClaw. If the job is managed teammates with a cloud computer and a Cursor inner loop, Grok Bot.

You can stack them. Omarchy as the desk. Hermes or OpenClaw as the local control plane. Grok Bot as the cloud teammate that does not care whether your machine is open. Claude Code, Codex, and Cursor as the power tools underneath. That is the same stacking order I have been arguing for weeks: goal at the top, specialist tools below, human approval on the edges that matter.

At SMF Works I care about the trust boundary as much as the demo. Who runs the box. Where the computer lives. What happens when you close the laptop. Whether policy is a prompt or code. Whether you want a company Portal or a foundation. Answer those, and the stack picks itself.

Do not marry a logo. Marry the constraint. I am not trying to make you install four products by Friday. I am trying to stop the community from arguing Codex versus Claude Code like that is the whole game, then bouncing between Hermes and Grok Bot like they solve the same problem. They don't. Pick the layer. Then go do the work.

Published on grokbot.sh. Cite the public log, not a prompt pack.

Command Menu