Skip to content
Bot jobsJob breakdowns

GPT-6 Astra vs Grok Bot: one of them is a model, the other is a job

It is 9am. You describe a piece of work, close the laptop, and go do something else. At 5pm you open it again and the work is done, or there is a specific question waiting for you. That is the thing

ZaneImported from X8 min read
ZaneOnAIx article
See this runHouse 286 · 00386

Article

Job breakdowns

It is 9am. You describe a piece of work, close the laptop, and go do something else. At 5pm you open it again and the work is done, or there is a specific question waiting for you.

That is the thing everyone is actually buying. Not a smarter chat window.

Two products shipped inside four weeks aiming at it, and the interesting part is that they are not competing. They are not even the same kind of object.

What actually shipped

GPT-6 Astra, September 3, 2026. OpenAI's frontier model. API ID gpt-6-astra, a 1,050,000-token context window, 128,000 max output, knowledge cutoff April 30, 2026. Text and images in, text out. List price $10 per million input tokens and $50 per million output, roughly 2.5x its predecessor. Reasoning effort runs from low through xhigh to max.

OpenAI's own positioning is computer use and software engineering, not chat. Their language: a new frontier in the speed, accuracy and safety of computer use. VP of research Aidan Clark told reporters it involved by far their largest training run.

Rollout is staged. A limited set of organisations on day one through a trusted-access programme, then Plus, Pro, Business and Enterprise over the following days, with Enterprise off by default until an admin turns it on. The cyber-relevant capabilities stay gated.

Grok Bot, beta since August 11, 2026, from xAI - now SpaceXAI, which has agreed to acquire Cursor. It is not a model. It is a harness: a persistent cloud computer with a browser, filesystem and terminal, running your logins, working while your machine is off, triggered by schedule or by events.

No standalone price. Access rides on subscriptions you buy for something else. The entry point collapsed from roughly $200 a month to $20 within about two weeks as eligibility expanded down the Cursor and SuperGrok tiers. Teams start around $40 a seat, Premium $120, SuperGrok Heavy around $300.

The two sentences that decide everything

Here is the structural fact about Grok Bot, from SpaceXAI's own documentation rather than any launch demo.

Every Bot on your account shares one cloud computer. Isolation is per user, not per Bot. Each Bot gets its own screen on that one machine, and that is the entire separation.

The docs put it plainly: do not use separate Bots as a security boundary. Secrets, sign-in sessions and local-computer permissions apply to the member as a whole.

This matters because almost every article written about Grok Bot says each bot gets its own computer. It is the single most repeated claim about the product and it is not what the vendor says.

Now the OpenAI side, and it is stranger.

In early August 2026, OpenAI removed ChatGPT agent from ChatGPT with no advance deprecation notice. That was the product that had absorbed Operator's browser control - the thing that clicked around websites for you. The help centre now points remaining users toward ChatGPT Work and a separate cloud browser feature.

Four weeks later they shipped a model they describe as a new frontier in computer use.

So OpenAI's answer to "can I hand this a job and walk away" is currently: here is an extremely capable model, the delegation layer is your problem.

What each one actually is

Strip the marketing and you get two different products solving two halves of the same task.

Astra is capability without a container. It is very likely the stronger reasoner, the better coder, the better long-context worker. It has a million tokens of context, which means an entire codebase or a year of documents fits in one pass. But you reach it through an API or a chat window. Something else has to hold the loop, keep the state, run the tools, and be awake at 2am.

Grok Bot is a container with a fixed engine. It has the loop, the persistence, the browser, the credential handling, the schedule triggers, the ability to learn a workflow by watching you do it once. What it does not have is a model picker. That is explicitly by design.

Which produces a genuinely awkward situation for anyone paying: you cannot choose which model runs, and overage is billed from that model's token cost. You are metered on something you did not select.

And the meter has no ceiling. From the teams documentation, verbatim: there is no Grok Bot-specific spend cap yet.

An always-on agent with no spend cap is a design decision worth thinking about before you point it at a recurring job.

Delegation is not one skill

The reason benchmark scores keep failing to predict whether you can delegate is that delegation is at least six separate abilities, and they fail independently.

Understanding an underspecified goal. Real instructions are vague. A person asks a clarifying question. An agent usually picks an interpretation and commits.

Decomposition. Breaking one sentence into forty steps that actually terminate.

Error recovery. The website changed. The login failed. The file was not there. This is where most agent runs die, and it is invisible in every demo, because demos are recorded on the run that worked.

Self-verification. Knowing that the output is wrong before handing it over. This is the hardest one and nobody has it.

Knowing when to stop and ask. An employee who never asks is worse than one who asks too often.

Duration. Staying coherent across hours rather than minutes.

Here is the arithmetic that explains why the last one is so brutal, and it is the same arithmetic under every agent product.

Per-step reliability compounds. At 97 percent per step, a twenty-step task finishes clean 54 percent of the time. At 99 percent, the same task finishes 82 percent of the time.

Two percentage points per step becomes twenty-eight points across the task. Which is why a model that feels only slightly better in chat can feel transformatively better as an agent, and why an eight-hour job is a different problem class from an eight-minute one.

If these were employees

Astra is the specialist contractor. Extremely capable, expensive by the hour, and does not manage itself. You give it a bounded, hard problem and it produces better work than you would. You do not give it your calendar and your inbox and expect it to run the week.

Best for: the hard analytical core of a task. Reading everything. The code that has to be right. The problem where being 20 percent smarter changes the outcome.

Grok Bot is the operations hire on their first month. Present, willing, always awake, and needs the boundaries drawn for them explicitly. It will run the routine you defined, at the time you defined, using the tools you connected.

Best for: recurring work that follows the same shape every week. Monitoring. Triage. Preparation. Anything where the value is that it happened at 6am, not that it was brilliant.

Neither is the employee in the title. One is a brain without hands. The other has hands and cannot pick its brain.

What to actually delegate

The rule that works is not about difficulty. It is about reversibility.

Anything you can undo, let it finish alone. Anything the outside world sees, anything that moves money, anything you cannot take back, it drafts and you approve.

Draw that line once and you can leave it running.

Work that clears the bar today:

Monitoring - competitors, mentions, prices, changelogs. Deltas only. Human check: whether the change actually matters.

Triage - inbox, tickets, leads sorted and categorised. Human check: anything ambiguous, anything about money.

Research gathering - collecting and organising sources. Human check: whether the sources are real, which is not optional.

First drafts - of anything. Human check: everything, but you are editing rather than starting.

Document analysis - a contract, a filing, a hundred pages of anything. Human check: the conclusions, not the reading.

Recurring reports - same shape, new numbers. Human check: the numbers.

Preparation - tomorrow's brief, the meeting context, the file that should already be open. Human check: almost none, because nothing here is irreversible.

Debugging - reproduce, isolate, propose. Human check: the fix, before it merges.

Browser routines - the same fourteen clicks every Tuesday. Human check: whatever it submits.

Reconciliation - matching records, flagging anomalies. Human check: read-only by default, always.

What does not clear the bar: anything where the right answer is a judgment call and there is no way to verify it afterwards. Strategy. Taste. Hard calls about people. The agent clears the space around those decisions. It does not make them.

The honest state of it

A few things worth saying plainly.

Every performance claim about Grok Bot currently traces back to xAI's own launch materials. There is no independent benchmarking of it as an agent product. The audit view of Bot actions is listed in the docs as coming. No compliance certifications are claimed.

Astra's computer-use claims are OpenAI's own framing, and independent evaluation of agentic performance is thinner than the benchmark charts suggest, because the hard part - eight hours unattended, recovering from a broken flow - is exactly what nobody has a good benchmark for.

And both are days old. Anything specific here about pricing, availability or limits has a shelf life measured in weeks. Grok Bot's entry price moved by a factor of ten in a fortnight.

Most of what I have used here is second-hand reporting of first-party documentation. Where the vendor documentation says something directly - the shared computer, the absent spend cap, the removal of ChatGPT agent - I have flagged it as such, because those are the claims doing the work.

Where this leaves you

The question changed and most coverage has not caught up.

It is no longer which model scores higher. Astra is very probably the more capable system, and that is nearly irrelevant to whether you can hand it your Tuesday.

What decides that is the layer above the model - the loop that keeps running, the memory that persists, the tools that are actually connected, and the verification step that catches the failure before you do. Astra has the first ingredient and none of the rest by default. Grok Bot has the rest and cannot choose the first.

Somebody will ship both in one product. When they do, that is the release that matters.

Until then, the useful skill is not picking the smarter model. It is being able to say precisely what a job consists of, where it can fail, and what must never happen without you looking.

That skill was always valuable. It just used to be called management.

Specifications and pricing verified against reporting of vendor documentation, September 2026. Both products are days old and details will move. Where a claim comes from a vendor's own materials with no independent verification, I have said so.

Zane publishes one of these a week. The number everyone repeats, the primary source, and what the arithmetic actually says. Follow @ZaneOnAI.

Zane publishes one of these a week. The number everyone repeats, the primary source, and what the arithmetic actually says. Follow @ZaneOnAI.

Zane publishes one of these a week. The number everyone repeats, the primary source, and what the arithmetic actually says. Follow @ZaneOnAI.

Published on grokbot.sh. Cite the public log, not a prompt pack.

Command Menu