Skip to content
Bot jobsJob breakdowns

A coworker who never sleeps: the complete Grok Bot guide

xAI shipped Grok Bot on August 11, 2026. Five Bots, one shared machine, $120 per seat. Here is the full build order, from the first hire to the two things that will quietly break it. Most people open

The maestroImported from X10 min read
maestroslayx article
See this runHouse 165 · 00210

Article

Job breakdowns

xAI shipped Grok Bot on August 11, 2026. Five Bots, one shared machine, $120 per seat. Here is the full build order, from the first hire to the two things that will quietly break it.

Most people open Grok Bot, create six teammates in ten minutes, name them after the jobs they hate, and end the week with six half-finished threads. The product punishes that approach because it is not a chat window with tools attached. It is an app where a named agent gets a real computer in the cloud, signs into the software already being paid for, and drives that software through the interface the way a person does. Treat it like a hire and it works. Treat it like a prompt box and it produces confident garbage on a schedule.

This guide runs the build in order: what the product actually is, the machine underneath it, the first hire, the roster, the demonstration, the review cycle, the cost, and the places where unattended work goes wrong.


First, know which xAI product this is

Three xAI products get confused with each other and the mix-up leads to the wrong purchase. Grok Build is a coding agent that runs in the terminal, released to beta on May 27, 2026. Build Mode is the consumer feature inside the Grok app that turns a description into a working app, game, or dashboard live in the chat, opened to every plan on web and mobile. Grok Bot is neither. It is a standalone app for persistent teammates who work across ordinary business software, launched in early beta on August 11, 2026 for desktop, including a posted Linux build, and iOS.

The distinction matters because teams evaluating Grok Bot as a developer framework go looking for APIs that do not exist. There is no SDK, no orchestration layer, no fleet of programmable agents. There are named coworkers a person messages from a phone or a laptop, and conversations pick up across either surface. Anyone expecting a swarm framework will be disappointed. Anyone expecting a chat assistant will underestimate what a persistent computer changes.


Start with the desk, not the hire

Every Bot on an account works on the same persistent cloud computer. This is the single most important sentence in the product and it appears in the documentation rather than on the launch page, which promises Bots with their own computer. The docs then repeat the warning twice: do not use separate Bots as a security boundary. Every signed-in session, every downloaded file, and every credential sitting on that machine is reachable by every Bot on the account. A Bot handling receipts is standing next to whatever session another Bot left open. Plan the account before planning the roster, because the account is the security model and nothing inside it is.

The practical version: one account per trust level. Personal admin lives on one account. Anything touching client data, money, or production systems lives on a second one with its own email and its own logins. That costs an extra seat and removes the failure that ends most agent projects before month two.


Hire one Bot and give it the worst hour of the week

The first hire should be the task that is boring, repeated, and easy to check. Something that takes 40 to 70 minutes, happens at least three times a week, and has a visible right answer. Inbox triage with a labeled output. A weekly report assembled from three dashboards. Invoice extraction into a sheet. The point is not ambition, it is verification. A Bot that produced the wrong summary of a strategy document is hard to catch. A Bot that pulled the wrong number off an invoice is caught in four seconds.

xAI's own teams started the same way. Their internal examples are a sales Bot that researches accounts overnight, scores contacts, and leaves drafted emails for approval, a demo-readiness Bot that checks the environment overnight and drops a checklist before morning calls, an ops Bot that processes invoices arriving in Gmail and seats new hires, and an engineering Bot that reproduces a bug in the product interface, files the ticket, and hands the fix to a second Bot for debugging. Every one of those has a checkable output and a human approval step in the middle.


Write the job description before the first message

A Bot performs at the level of the brief it was given, and the useful brief looks nothing like a prompt. It reads like an onboarding document for a new hire on day one. Where the inputs live, what tools to open, what the finished output looks like, where it gets saved, what to do when something is missing, and what never to touch. The last two lines carry the weight. An agent without an explicit escalation rule will invent a plausible answer instead of asking, and an agent without an explicit boundary will wander into the next open tab because it seemed relevant.

Name the Bot after the function rather than the mood. A roster of five named workers with five written briefs is a system a person can audit in a morning. A roster of five that were spun up conversationally is a system nobody can reconstruct three weeks later.


The roster: five jobs that survive contact with reality

MARA, inbox and receipts. Reads the mail, pulls invoices and receipts, drops them into a sheet with vendor, amount, and date, and flags anything that does not match a known vendor. Output lands in one file, which makes the morning audit take under two minutes.

COLE, calls into records. Takes transcripts and voice notes from the day before, writes them into the CRM or the notes system in a consistent format, and marks deals that stalled. The value is not the writing. It is that the record exists at all on days when nobody feels like writing it.

RINA, overnight research. Scans named sources against a fixed list of companies, tickers, or topics, and delivers one brief before the morning starts. Real limits matter here. A brief on 40 items with sources beats a brief on 400 without them.

VINCE, quality checks. Runs the same click-path through a product or a dashboard every night, screenshots what changed, and files a ticket when something is broken. This is the job computer use was made for, because it needs no API and no integration.

OWEN, admin closing. Handles the residue: scheduling, form filling, seat provisioning, the tasks that live in six systems with no connection between them.

Five names, five outputs, five files. Bots can message each other, share threads, and coordinate in group chats, so COLE handing a stalled deal to RINA for research is a supported pattern rather than a hack. Build those handoffs after each Bot works alone, not before, because a broken handoff between two unreliable Bots is impossible to debug.


Show the job once, and show it correctly

Grok Bot learns by demonstration. Walk a Bot through a task one time and it saves the path as a routine it can run on a schedule. This is the fastest onboarding in the category and it carries an obvious trap: a demonstrated routine inherits every mistake in the demonstration. A wrong filter, a skipped verification step, a tab opened out of order, all of it gets baked in and repeated at 3am for a month. Slow down for the single pass that teaches the routine. That one pass is worth more than any prompt engineering applied afterward.

Two habits pay for themselves. Demonstrate on a real workload rather than a clean example, so the Bot sees the messy version it will actually meet. And demonstrate the failure case at least once, showing the Bot what to do when the file is missing or the number does not reconcile, so the routine has somewhere to go other than forward.


Design the approval, not just the task

The whole pitch of the product is work that comes back finished, with the Bot returning only when something needs a human call. That return point is a design decision and leaving it to the model is how teams end up either approving everything or approving nothing. Pick the trigger in advance: any amount over a threshold, any new vendor, any message going to a person outside the company, any change to a live system. Everything under the line gets done and logged. Everything over it waits in a queue.

A well-tuned roster produces a short morning queue and a long list of completed items. A badly tuned one produces either a Bot that pings constantly, which is worse than doing the task manually, or a Bot that never pings, which is worse than that.


The first scheduled run is the real interview

Nobody watches the first unattended run and almost everybody should. A Bot that performed perfectly under supervision behaves differently when a page loads slowly, a login expires, or a modal appears that was not there during the demonstration. Sit through the first scheduled execution, check the second, then move to spot checks on a fixed day each week.

There is a security reason as well as a quality one. The OWASP GenAI Security Project spent 2026 ranking agent goal hijacking as the highest-priority agentic risk, and a Bot reading untrusted content on a shared machine is exactly the path that ranking describes. An email, a document, or a web page can carry instructions the Bot was never meant to follow. Security vendors put the share of enterprise agents running without any oversight around 47%, a figure worth treating as vendor-reported and directionally true.


Know where it is slow

Driving an interface at human speed costs real time. Where a proper connector or MCP server exists, the Bot moves at API speed and finishes in a fraction of the time. Where it does not, the Bot clicks like a person with unlimited patience. Both are useful and only one is fast. Route the high-volume work through connectors, and save computer use for the systems that have no other door. Grok has supported custom MCP servers on paid accounts since April 3, 2026, so the connector path is worth checking before assuming the click path is the only option.

This also shapes what to schedule when. Slow computer-use jobs belong overnight, where the wall-clock cost is invisible. Fast connector jobs can run hourly during the day without anyone noticing.


What it costs and where the price lives

Grok Bot is not sold on an xAI pricing page. Access runs through Cursor at $200 a month on Cursor Ultra and $120 per seat a month on Cursor Premium Teams, and it comes bundled for anyone already paying for SuperGrok Heavy or Cursor Ultra. Coverage widened later in August to SuperGrok Plus and Cursor Pro Plus, with enterprise access on a waitlist. Authentication, single sign-on, and privacy settings all run on Cursor accounts, which means an admin change on the Cursor side propagates to the Bots.

One clause stops regulated teams cold: Grok Bot requires data storage and does not support Legacy Privacy Mode, so an admin has to change that setting before the product opens at all. Teams under strict data policies should settle that question before anyone builds a roster.


The one category to keep on a short leash

Anything with withdrawal rights stays off the Bot account. Read-only is the ceiling for banking, brokerage, and payment systems, and it should arrive through a scoped connector rather than a stored password. The reason is the shared machine again. A Bot doing overnight research and a Bot with a live session on a money system are the same Bot in practice. Let the roster build the case, prepare the file, and draft the message. Let a human press send on anything that moves value. Regulated firms have an extra reason, since supervisors now expect a clear answer to which process took an action and under what authority, and a shared login cannot produce that answer.


A 30-day rollout that does not collapse

Week one: one account, one Bot, one task, fully supervised, run daily and checked every time. Week two: put that task on a schedule, watch the first two unattended runs, and write down every deviation. Week three: add the second and third Bots on unrelated tasks, keeping their outputs in separate files so failures stay isolated. Week four: connect two Bots through a handoff, set the approval thresholds, and do a full audit of what the account can reach.

The audit is the step people skip and the one that determines whether this scales. List every tool the account is signed into, remove the ones no Bot needs, and confirm nothing on the machine has spend or withdrawal rights. Do that before the roster grows past three.


What sits under the hood, and what does not

xAI has never published which model runs Grok Bot, so nobody can benchmark the product against a version number. The lineup around it is documented and moving fast. Grok 4.5 arrived on July 8, 2026 at 1.5 trillion parameters with a 500K context window and a first-place finish on agentic tool use in Artificial Analysis testing. Grok 4.6 followed in August, ranked third overall, and landed in Google's Model Garden and GitHub Copilot. The teammates stay a black box while the models beside them ship spec sheets, which is a fair trade for a beta and a poor one for anything that has to pass a vendor review.


When not to build this at all

Some work does not belong to an unattended agent yet. Judgment calls with no checkable output. Anything where a wrong answer is expensive and slow to detect. Client-facing communication with no approval step. One-off tasks that will never repeat, where the demonstration costs more than the job. And any system where a mistake cannot be reversed. The roster earns its money on volume, repetition, and reversibility. Give it work that has all three.

Run it properly for a month and the uptime counter becomes the pitch on its own. Four hundred hours of work nobody had to be awake for, handed back one file at a time. The catch is that all five of them are sitting at the same desk, and the desk was never locked.

Published on grokbot.sh. Cite the public log, not a prompt pack.

Command Menu