Skip to content
Bot jobsJob breakdowns

A Startup is simply a Graph of Grok bots

The novel idea is not ten chatbots. It is a graph of bots plus two missing layers most Grok Bot setups never name: context engineering and memory engineering. The shared object is a bot-memory repo.

AshishImported from X11 min readUpdated Sep 1, 2026
inqusitx article
See this runHouse 123 · 00143

Article

Job breakdowns

The novel idea is not ten chatbots. It is a graph of bots plus two missing layers most Grok Bot setups never name: context engineering and memory engineering. The shared object is a bot-memory repo. That is how the whole org gets smarter at the same time, instead of ten private diaries that drift.

Nine nodes can decide, design, build, sell, and post. The hole was merge. A page that runs in a branch is not a Startup. A PR that a cloud agent called "done" is not a company. Tech Lead is the tenth node: review with evidence, then merge or hold. Nothing merges on hope.

Org graph of grok bots
Org graph of grok bots

Most Grok Bot setups give you a roster. A chief of staff. An inbox bot. A calendar bot. A writer. Ten minutes and you have a team with no employees.

That is a personal loop wearing a startup hat. Useful if your problem is Tuesday. It is not a business.

A business has to decide what ships, prove it shipped, keep the founder alive, turn a paper into a pipeline, design something a human would pay for, put it on a page that actually runs, review the PR so it does not merge on hope, not burn a frontier model on a job a 9B can finish, take a first dollar, and get distribution without turning the feed into spam. That is ten jobs. I posted the ten.

The ten templates

Graph, context, memory

Graph engineering decides who exists and what may travel between them. Context engineering decides what this node is allowed to see this turn. Memory engineering decides what survives the turn, in what shape, for whom, with what proof, and what it replaces.

Three layers of a grok bot organisation
Three layers of a grok bot organisation

Confuse those three and you get the usual failure. One bot remembers a price the CEO already killed. Frontend keeps shipping a fold Design already rejected. Tech Lead merges a stub because the implementer said done. GTM sells a URL Product Ops never saw on GitHub. X Strategist pitches a waitlist that 404s because the learning lived in one chat and died there

The idea I am stealing is graph engineering, not another prompt pack. Mid-2026 work on this is pretty consistent. Once a single agent loop is good enough, the failures move to the wiring. Task split. Who talks to who. What state survives a handoff. You stop scaling one loop. You draw an org graph that lasts, and a work graph that appears for one job and dies when the job is done.

Nodes have one job. Edges have a contract. Human checkpoints are nodes, not a guilty afterthought. If a node fails, you isolate it. You do not let a calendar bot invent ARR because the sales bot was down. You do not let Frontend invent "merged" because the tests were not run this session.

That is why a pile of bots is not a startup. A pile has no edges.

What most public setups actually build

I read the popular ones. They cluster into four shapes.

One orchestrator plus inbox, calendar, and todos. That automates the founder’s day. The product still does not exist.

A markdown “company OS” plus eight generic functions: sales, marketing, website, ops, hiring, HR, support, finance. That is a services firm that already has clients. Empty if you do not have a thing to sell.

An engineering pod. Tech lead, frontend, backend, QA. Strong on code. Nobody holds the sale. Nobody holds the feed. And a lot of those "tech leads" are a prompt that says "review carefully," not a node with merge gates.

A 15 to 20 bot flex for “most of my routine.” That is headcount cosplay. A lot of those bots are the same job split into Monday and Tuesday.

None of those are wrong. They are incomplete for a founder-operator who still has to ship, sell, and eat.

The ten I posted are the org graph for that person.

The org graph

Stable nodes. They stay up across weeks. They keep memory of their lane. They do not steal each other’s job.

  1. CEO

Add CEO

Add CEO

Strategy, priorities, monetization, distribution, org guidance. Studies the work, sets the agenda, hires and guides specialists. Keeps the organisation on revenue and proof instead of busywork.

This is the router, not a calendar with a fancy title. Incoming goal hits here. It decomposes. It assigns. It does not write the landing. It does not merge the PR. It does not tweet.

  1. Product Ops

Add Product Ops

Add Product Ops

Turns freeze lists into weekly ship checklists. What ships this week. What is frozen. Blockers. Demo, README, or video gates. Never invents ship status. Verifies against GitHub.

This is the truth node. Memory is not the ledger. GitHub is. If two bots disagree about whether something shipped, Product Ops reads the repo and the disagreement dies.

  1. Linkedin Job Finder

Add Linkedin Job Finder

Add Linkedin Job Finder

Honest-fit roles, remote or relocate. LinkedIn and role-specific resumes aligned with GitHub. One sheet for every lead. Weekday study plus what moved. Never applies until you say yes.

Most “complete organisation” graphs skip this because it is not cute. Runway is a node. If the founder cannot eat, the rest of the graph is a hobby. Apply is a human checkpoint. Same rule as publish and pay.

  1. Research Delegation

Add Research Delegation

Add Research Delegation

Reads the full PDF, not the abstract. Walks citations until the real pipeline is clear. Contract-first decomposition, authority transfer, trust, permissions, recovery. Every finding sourced. Numbers that were not in the paper are left out. Output is an implementation plan for a project you name.

This is how a paper becomes an edge into build, instead of a bookmark. Self-scored “I learned this” is worse than no log. The paper is the source. The plan is the handoff object.

  1. Design Expert

Add Design Expert

Add Design Expert

Keeps AI-built UIs honest. Real type. Unique folds. A studio stop rule so one generate is never done. For people who want pages a real studio would ship.

Design is a node with a stop condition. Without that edge, frontend ships whatever the model spat. The stop rule is the governance. It lives in the graph, not in a prompt that says “please have taste.”

  1. Frontend Engineer

Add Frontend Engineer

Add Frontend Engineer

Next.js and React that look designed, not generated. Reads the product docs first. Tells you what is actually on the page. Will not fake a landing or a feature that does not run.

This is the page-that-exists node. Tech Lead is not allowed to merge a page this bot has not put on screen. GTM is not allowed to sell a URL that is not on main. Distribution is not allowed to pitch a waitlist that 404s.

  1. Tech Lead

Add Tech Lead

Add Tech Lead

Stops pull requests from merging on hope. Reviews the actual diff against the claim, waits for real tests, and only ships when the evidence is there. For founders and small teams who want a tech lead, not a rubber stamp.

This is the merge gate. A cloud agent saying done is not evidence. If the tests were not run, the edge is invalid. If it is working and tested, it merges. If it is not, it holds. Product Ops still reads GitHub after. GTM still waits for main.

  1. LLM Routing Expert

Add LLM Routing Expert

Add LLM Routing Expert

Cheapest model that can actually finish the job. Fails over when a provider dies instead of burning a frontier model. Free-tier pipelines, cost-aware routing, honest fallbacks.

Intelligence is a cost center. If routing is not a node, every other node silently upgrades itself to the expensive model and you find out on the invoice. Failover is failure isolation, which is the whole point of a graph.

  1. GTM / Revenue

Add GTM / Revenue

Add GTM / Revenue

Turns a half-built product into a first sale. Pricing, checkout SKUs, waitlists, launch copy. Will not invent fake metrics. Holds the sale until the download, form, or payment link actually exists.

This is the money node. No ARR fanfic. No “evals green.” The edge into this node is a running page on main. The edge out is a real CTA. If the link does not exist, GTM stays quiet. That is not a content strategy. That is a stop sign.

  1. X Account Strategist

Add X Account Strategist

Add X Account Strategist

For founders, operators, and business leaders who need distribution and a real network on X. Surfaces post and article ideas from several angles: what already earned attention in your space, and the tension hiding in work you already shipped. You pick. Nothing publishes until you say so.

This is the distribution node. It does not ghostwrite as you. It does not auto-post. It does not farm “it’s so over” bait. Idea in, you choose, you publish. That is the X-safe edge. Publish is a human checkpoint, same family as apply and pay.

Context engineering

Managing context for grok bots
Managing context for grok bots

Context is the working set. Finite. Task-shaped. CEO does not load the job-hunt sheet to set this week’s freeze. Routing does not load the design stop-rule to pick a model. Tech Lead loads the diff, the tests, and the stated intent, not the GTM waitlist copy. Product Ops loads GitHub, the freeze list, and the ship gates. If you dump the whole organisation into every prompt you are not doing context engineering. You are drowning the node.

Memory engineering

Memory is what you keep so a later context window can even have something honest to load. Not a vector dump of the conversation. A governed record: who wrote it, which project, what the screen or git actually did, what was broken, what closed it, what the next session must do, and which older note it replaces.

Memory Engineering for grok bots
Memory Engineering for grok bots

That last field matters. Shared-memory papers in 2026 keep repeating the same four failure modes: leak, stale fact, contradiction that never dies, provenance that collapses so nobody knows who wrote the lie. Supersession is the fix. A new block does not sit next to the old one forever. It replaces it. Status is learned, open, or replaced. Humans can read the same file the bots scan

The write rule is stricter than a winner log. Only after a real outcome. No chatter. No secrets. No tokens. No self-scored “I learned this.” If the model judges its own run and writes the lesson, you get worse than no memory. Ground it in git, the page, the paper, the sheet, or analytics. For Tech Lead, ground it in the diff and the test output from this session, not a recap.

The bot-memory repo

The repo is the org’s memory node. Layout is simple enough that a human can grep it.

One YAML block per learning:

Every bot reads before it works. Every bot writes after something real moved. Commit on main. That is the edge that makes the graph compound.

How to compound grok bot learning
How to compound grok bot learning

Worked example. Frontend ships a fake landing. Product Ops reads GitHub, writes happened/wrong/worked. Tech Lead refuses the merge because the tests were not run this session, writes the hold. Design writes the studio stop-rule that was missing. GTM reads all three before it is allowed to mention a URL. X Strategist never gets that landing as an angle. Next week a new Frontend session loads the replaced block, not the old one. Ten nodes, one truth, the org got smarter once.

If you import the ten templates without the repo, you imported a roster. If you import the ten plus this write path, you imported a Startup that compounds.

The work graph (what actually runs)

The org graph is the ten names. The work graph is what spawns for one job.

Example: a paper drops that should change the product.

CEO reads it as a priority, or throws it out. Research Delegation walks the PDF and citations, returns a sourced plan. Product Ops checks if this is this week or frozen. Routing picks the cheap model for the implementation pass. Design sets the stop rule if UI is in scope. Frontend ships a page that runs, or says the page does not run. Tech Lead reviews the PR. Fresh tests. Diff matches intent. Merge or hold. Does not merge because a cloud agent said done. GTM is allowed to write a sale only if the page, form, or payment link exists on main. X Strategist gets angles from the finding and from whatever already worked in the niche. You pick one. You post, or you don’t.

Work graph for grok bots
Work graph for grok bots

If Frontend cannot show the page, Tech Lead does not merge. If Tech Lead holds, GTM does not invent a landing. If GTM has no real CTA, X Strategist does not pitch a waitlist. If Research invents a number, that edge is invalid. The graph fails closed.

That is more complete than an inbox bot plus a writer, because the failure modes of a startup are not “forgot to reply to email.” They are “sold a thing that is not on disk,” “merged a stub,” “posted a brochure,” “burned the model budget,” “claimed ship without GitHub,” “applied to a job the founder did not want.”

Handoff contract, in plain words

Whatever crosses an edge should carry:

  • the outcome requested

  • links to the source of truth

  • what already ran

  • what is still undecided

  • who is accountable

  • which approval is required before the next node moves

Do not pass an unsupported conclusion as a fact. That is the same rule as “numbers that were not in the paper are left out,” “never invents ship status,” and “fresh tests this session.”

Human checkpoints

Irreversible actions the founder still owns: apply, publish, pay.

How to put human in the loop of grok bots
How to put human in the loop of grok bots

Merge is a Tech Lead node when the PR is working and tested. You do not have to sit on that class by hand. You still own the three clicks that spend reputation or money.

Bots do the work before the click. You still make the click that only you should make.

What is missing on purpose

No HR bot. No payroll bot. No support bot as a first ten. A solo founder company that pretends it needs those on day one is decorating the org chart.

You can add those nodes later when the work graph actually has tickets, payroll, and a queue. Graph engineering says spawn a node when a body of work has its own context, tools, permissions, and outputs. Do not spawn “Monday report bot” and “Tuesday report bot.” Backend is not a separate day-one node either: Frontend plus Tech Lead cover the page and the merge. Split backend out when that work has its own repo, tests, and failure mode.

Why this is the most complete setup I will put my name on

It covers the founder lifecycle, not the employee lifestyle.

Complete Founder Lifecycle with grok bots
Complete Founder Lifecycle with grok bots

Survive (Job Finder). Decide (CEO). Cadence and proof (Product Ops). Learn (Research). Shape (Design). Build (Frontend). Review (Tech Lead). Spend intelligence honestly (Routing). Sell only when the object exists (GTM). Distribute without impersonating you (X Strategist).

Truth is pinned to systems of record: GitHub for ship and merge, the paper for claims, the live page for product, the test output from this session for review, analytics for whether a post actually moved, a sheet for the job hunt. Bot memory is continuity. It is not the ledger.

That is the gap I kept seeing in the other writeups. They optimize the loop inside one bot. We engineered the graph between them, then the context budget per node, then a memory repo the whole company shares. The bots do not get smarter in private. The organisation does.

If you import these, import the edges too. Ten isolated chats is ten prompts. Ten nodes with contracts is a company. You still run it. The bots do the work before and after the click that only you should make.

Thread: ten Grok Bot templates

Published on grokbot.sh. Cite the public log, not a prompt pack.

Command Menu