What I got from 3 days of Grok Bot Galaxy
Grok @Bot is @xAI's agent product. You set up persistent bots, each with its own instructions and memory, and they run on a cloud computer with a browser. They can work on a schedule while your laptop
Article
Job breakdowns

Grok @Bot is @xAI's agent product. You set up persistent bots, each with its own instructions and memory, and they run on a cloud computer with a browser. They can work on a schedule while your laptop is closed. This week xAI ran Grok Bot Galaxy, a 3-day livestream from San Francisco. Three hosts from the @Grok Bot team built a business on camera from nothing, and between build sessions there were guest talks and workshops.
I run Silkster, a studio that builds and manages websites for businesses and mission-driven organizations. We also set up automations that take work off their plate. Ten days before the stream started, I'd set up eight Grok Bots for the parts of the business that live behind a browser. Watchman checks every client site each morning. Prospector looks for local businesses whose websites are costing them work. Docent answers questions about how Grok Bot behaves on my account, so I don't have to guess what the platform can do.
That meant I watched with a specific question. What should change in a fleet that's already running?
Here's what I took from each day, and what's live in my fleet now because of it.
Day 1: make every bot prove it
The first day was mostly about how the hosts set up their own bots. The idea that stuck was proof-gated completion. A bot can't say "done" without evidence, like a screenshot for a UI change or numbers for a performance fix. If the proof is missing, another bot asks for it before a person ever looks.
Codie Sanchez, who has bought more than a thousand small businesses, made the same point from the business side. People don't take claims on faith. So keep proof of everything you build, starting now, because in six months you'll want it and won't be able to recreate it.
Two smaller Day 1 ideas ended up mattering to me:
-
Define "urgent" once. Tell an agent something is urgent and it guesses what you meant. In the engineering workshop, the presenter defined P0 once, in plain language, and pointed every bot at that definition.
-
Let a scheduled run say nothing. Scheduled bots turn into noise fast. The fix is a written rule: if you found nothing, say so in one line and stop.
Day 2: the biggest fleet on stage came with a warning
Day 2 was workshops for sales, support and engineering teams. Simon, a presenter running one of the largest setups shown, closed by admitting he'd had way too many bots at times. His words were that it got "more chaotic and it does more harm than good."
That shaped my whole expansion. By the end of the stream I had thirteen candidate bots on my list. Four made it, and one of those four is a bot I'd first rejected.
Day 2 also gave me the most useful prompt of the event. Lauren, the host running engineering, asked her fleet to zoom out and review every bot against her goals. Then she asked where the bottleneck was. The answer came back that she was. Every change was still waiting on her to merge it by hand. I've put that question into a monthly audit, including permission to name me as the bottleneck.
Day 3: 300 pull requests and a pile of launch bugs
On Day 3 the hosts launched what they'd been building, which by then was an online card game, to real players. Most merges were automatic at that point. Over the three days the team merged more than 300 pull requests, most of them after a QA bot played through the game. Lauren said on stream that she hadn't looked at the code at all.
Launch day found the bugs. Matchmaking paired every player with the computer. Match history showed the wrong data. The advertising feature was simply broken. Hosts and players caught those by looking at the screen, and the QA bot missed them.
Once they were in production they added one rule for every fix. Reproduce the bug before writing any code. I'm adding that rule to how I handle client bug reports.
What's live in my fleet now
Groups, as a brake on adding bots. The business bots sit in three groups, and each group has a rule a bot has to meet to join. Watch bots read sites I'm paid to keep working. Market bots read the outside world and never contact anyone. Build bots touch the platform and run only when I trigger them. A fourth group holds one personal bot that stays separate from the business. The groups are labels in the app and don't enforce anything. They make me justify every new bot. If a proposed bot doesn't fit a group, it needs an argument for a whole new group.
One set of reporting rules, pasted into every bot:
-
Never write to Notion, where my tickets live. Only one thing writes there, and it isn't a bot.
-
If a run found nothing, report one line.
-
Attach evidence to every claim. A claim with no URL or screenshot behind it gets labeled as a guess.
-
Never report a clean run you didn't do. "Nothing's wrong" and "I couldn't look" can't read the same way.
-
P0 means a site is down, a certificate has expired, or a site is serving someone else's content. Nothing else counts.
Upgrades to four existing bots:
-
Watchman now walks every page on every client site. It checks each one on desktop, then again at phone width. It also saves a dated screenshot of each site on every run. That's the proof library Codie was talking about, and it builds itself every morning.
-
Docent runs a monthly audit of the fleet. It reports which scheduled runs fired, which bots produced nothing worth reading, and where work is piling up (and on whom).
-
Prospector has a teardown mode. Give it a domain and it reports what's broken on that site, with evidence and no pitch. This came from Shubh, who ran the founders workshop. His sales bot finds a real bug on a prospect's site before every call. His favorite example was a cookie banner covering the submit button.
-
A research bot that finished its one big project now runs a small quarterly check to keep that data current.
Three new bots:
-
Bellwether watches competitors' changelogs, pricing pages and job posts. A bot shown on Day 1 signed up for competitor products using throwaway emails. Mine sticks to public pages. If something is behind a signup, it reports that and stops.
-
Gauge measures speed and accessibility on every client site once a month and tracks the trend. That trend is what a client's monthly care plan pays for.
-
Surveyor finds local businesses that already spend money on promotion. Silkster is building a local business directory, and those are the businesses likely to buy ad slots in it. The idea came from a Day 3 bot pitched as a complete go-to-market brain. When the presenter ran it live, it did something much narrower and found real advertisers for one product. I kept the part that got used and named the bot for it.
And Critic, the one I'd rejected. Critic reviews pull requests. I'd crossed it off my own list because I already review code with Claude, Anthropic's coding agent, outside Grok Bot. I came back to it because review is where my build pipeline slows down, and I wanted a reviewer sitting right next to the builder. The pipeline starts with Drafter, which writes a design. Claude reviews the design against my architecture, and I approve it. Then Wright builds it and opens a pull request. Wright can't merge, and GitHub branch protection enforces that. Critic reviews what Wright opens and reports to me. No bot checks its own work, and I'm still the one who merges.
Plus Dr. Eggbot, from the marketplace. Lauren built Dr. Eggbot and published it on the Grok Bot marketplace. The hosts used it as a bot factory: describe the bot you need and it drafts one. It sits in my Build group now. It wrote Critic's first spec by searching my own notes. It's also how I confirmed bots can hand work to each other. Dr. Eggbot sent Docent a question, and Docent answered back without me in the middle.
What Docent found that I didn't ask about
Before expanding anything, I had Docent answer four open questions about how the platform works on my account. Two things it found on the way mattered more than the answers:
-
Every bot on an account shares one computer. Same files, same browser logins. Separate bots feel like separate employees, but there's no wall between them. A login one bot makes is a login they all have. That's the reason none of my bots ever signs into a client's systems.
-
Plugins are account-wide. Connect Notion for one bot and every bot can reach it. A bot's instructions tell it what not to do, but nothing stops it. So the "never write to Notion" rule has a second layer. Docent's monthly audit looks for any Notion page a bot created or edited.
What I didn't build
My test for a new bot is four questions:
-
Does the work recur on a schedule?
-
Does it need a browser, where no API or connector would do?
-
Can it do the job without writing in my voice?
-
Does it avoid sending anything to someone outside the company?
A no on any of them means the work belongs somewhere else. Usually that's Claude, which has my files and my voice guide, or my ticketing system.
A lot of the bots recommended on stage fail that test. A voice bot trained on my sent mail. A bot that manages a knowledge base in Notion. A chief-of-staff bot that runs all the others. A bot that emails churned customers. I passed on every one of those, and the game's play-testing bot became a Claude skill for my site builds instead.
The honest parts
In three days, I never heard a real cost number. The closest was in the Grok Bot docs, which Docent pulled for me. Usage is metered by agent steps and tokens. The docs warn that a tight schedule can burn a week of usage in a day. So my bots run on daily, weekly or monthly schedules, and I'll measure Watchman's deeper walk after its first run instead of guessing.
The self-improving-fleet pitch also has a gap. On Days 2 and 3, the defects that mattered were caught by a person looking at the output. My bots run the checklist well. Deciding what goes on it is still on me.
Restraint was what I took away. My business fleet went from eight bots to thirteen over the stream, and the list of bots I decided not to build is almost as long.
Why a web studio runs any of this
Most of this fleet exists for one reason. Silkster designs, builds and manages websites for businesses and mission-driven organizations, and sets up automations for them. Once a site launches, somebody has to keep it working. Most owners already wear five hats. Checking their own website every morning shouldn't be a sixth.
That's the job of a Silkster care plan. Watchman walks every client site each morning, and Gauge tracks speed and accessibility month over month. When something breaks, it's on my list before most customers would ever notice. The owner gets a site that works. They don't need to know or care that bots are involved.
I also set up automations like these for other businesses. Usually it's one or two that take a real chore off your plate, like a weekly summary pulled from the tools you already use, or a daily check on something you're tired of checking by hand. I pick them the same way I picked mine. Each one has to earn its spot, and a person still signs off on anything that reaches a customer.
If your site is costing you work, or you've been putting off a rebuild for a couple of years, send me your domain. I'll send back what's broken on it, with evidence and no pitch attached. If there's a chore you'd love to hand off, tell me about it. Either way, you can decide from there if you want to talk. silkster.com
Published on grokbot.sh. Cite the public log, not a prompt pack.