Grok 4.7: The Next Reasoner
Elon Musk has now put a clock on it. Grok 4.7 comes out in ten days. That sentence is not a rumor mill item. It is a public schedule from the person shipping the model. Grok 4.6 is already in the
Article
Job breakdowns

Elon Musk has now put a clock on it. Grok 4.7 comes out in ten days.
That sentence is not a rumor mill item. It is a public schedule from the person shipping the model. Grok 4.6 is already in the field. It sits first on GPQA Diamond. It has taken first on CursorBench in extra-high thinking mode. It has tied the top of the Artificial Analysis Agentic Index. It has led MedAgentBench. It is live on Amazon Bedrock. It is the model inside Grok Build, Grok Bot, Grok Voice, and the Grok app.
4.7 is the larger pretrain.
Musk has described it as a 2.1 trillion parameter model, up from the 1.5 trillion parameter base behind 4.6. He has said it will be better than 4.6 in every way except that it will be slightly slower to serve, with even better token efficiency. He has also said initial training finished and that a massive amount of SpaceX company data went into supplemental training. The claim that follows from that corpus is specific: he would be shocked if any current model is better at real-world engineering.
That is the frame. Not a mystery box. A larger reasoner, trained into the same product surface, with an engineering corpus no consumer chatbot company can copy.
Start with what 4.6 already is, because 4.7 only matters as an upgrade to a working system.
GPQA Diamond is 198 graduate-level questions in physics, chemistry, and biology. They were written to resist casual search. Domain PhDs average about 65 percent. Skilled non-experts with the open web sit near 34 percent. Grok 4.6 (high) is at 94.9 percent on the Artificial Analysis board. That is first. Gemini 3.7 Flash sits next. GPT-5.6 Sol follows. Claude Fable 5.1 is on the same chart and still below. The gap at the top is tight. Holding first after new frontier releases is the part that tells you the post-training stack is not a one-week spike.
This benchmark is easy to dismiss if you only care about chat tone. It is harder to dismiss if you care whether a model can keep a physical system in working memory. Energy. Bonds. Forces. Rates. The next state of a process that does not announce itself in the prompt. That is the same muscle a civilization needs if it is going to put compute in orbit, land machines on other worlds, and keep those machines inside temperature and power limits.
CursorBench is the coding test Elon actually pushed people toward. He did not only quote a chart. He told people to try 4.6 in the Grok Build harness or the Cursor app. Extra-high thinking mode took first at a much lower cost per task than Fable 5 Max and Opus 5 Max. In the comparison that circulated with his engagement, Grok 4.6 Extra High sat at 70.8 percent and about $2.81 per task. Fable 5 Max was 70.5 percent at about $17.32. Opus 5 Max was 70.0 percent at about $8.23. GPT-5.6 Sol Max was 67.2 percent at about $5.69. First place at a fraction of the spend is the product thesis: intelligence that can stay on a long job without burning a ridiculous amount of compute.
The Agentic Index is the other board that matters. Tool use. Planning. Autonomy. Complex problem solving. 4.6 tying Claude Opus 5 at the top of that index is why Bot and Build are not just skins on a chat model. A score on a question set can be bought with fluency. An agentic score has to survive tool calls, state, and the ugly middle of a task.
MedAgentBench is the same idea in a clinical EHR environment. The model has to act, call tools, and complete multi-step work across hundreds of clinically derived tasks, not only answer a quiz. 4.6 took first at about 95.9 percent pass@1, ahead of GPT-5.6 Sol and ahead of Grok 4.5. Elon’s public line on that result was blunt: 4.6 is number one on healthcare questions. Healthcare is not a playground domain. A model that executes the wrong workflow is not charming. It is dangerous. The fact that two Grok generations sat in the top three is a family signal, not a one-off.
4.6 also set the economic baseline. On one Artificial Analysis comparison, it matched a 61 intelligence score against GPT-5.6 Sol while pricing at $2 input and $6 output per million tokens against $5 and $30. That is 2.5 times cheaper on input and 5 times cheaper on output for the same headline score on that particular board. Elon’s public line on that chart was simple: try 4.6, and 4.7 is a major upgrade coming soon.
So the incoming model is not arriving into empty air. It is arriving into an installed stack.
The parameter jump is the first physical fact.
1.5 trillion to 2.1 trillion is not a rounding error. A larger pretrain, if the data and the training recipe hold, buys more internal capacity for long chains of constraint. Physics problems. Multi-file codebases. Propellant and thermal budgets. Procedures that have a correct order and a wrong order. The cost of that capacity is latency. Musk already said the quiet part: slightly slower to serve. The compensating claim is token efficiency. Fewer wasted tokens per unit of useful work.
For a chatbot, token efficiency is a bill. For an agent, token efficiency is survival. A Bot that burns the budget before the form is submitted is not a teammate. It is a demo that died in the hallway. If 4.7 spends fewer tokens to reach a correct action, the product surface changes even if the first token arrives a little later.
If that trade is real, 4.7 becomes the model you point at the hard task and 4.6 remains the model you point at the fast loop. That is how product lines actually work. Not one model for every job. A family. Fast for the short hop. Dense for the long burn.
The SpaceX corpus is the second physical fact.
Most frontier models are trained on the public internet plus licensed libraries plus synthetic traces. That is a lot of language. It is not the same thing as the internal texture of a company that welds stainless tanks, fires 33 engines, flies a constellation, recovers a ship from the water off Christmas Island, and iterates a heat shield from samples taken in the field. Supplemental training on SpaceX company data is a bet that the model should know how real engineering looks when it is written down by the people doing it.
That does not make 4.7 a replacement for a propulsion engineer. It does mean the model should be less likely to hallucinate a clean textbook answer onto a messy vehicle problem. Real-world engineering is constraints, interfaces, failure modes, and the next test that would falsify the current design. If the corpus is as unique as claimed, the advantage will show up in those tasks first: structures, avionics procedures, manufacturing sequences, launch operations, constellation operations, not in a poetry bake-off.
Musk’s own line is the one to keep. He said 4.7 will exceed current models, that Anthropic is a great company and will release improved models, and that the SpaceX training corpus is so unique he would be shocked if any model is better at real-world engineering. That is a falsifiable claim. The day 4.7 is public, people can put the same engine, tank, thermal, and operations questions to every frontier model and see who stays inside the constraints.
What 4.7 should be able to do, if the public claims hold, is sit underneath the products people already use.
In the Grok app, it should be the model you put on the super-tough task. Elon has already said he still finds the app more useful than Bot for some work. A larger, more efficient reasoner makes that statement stronger if it can hold a long scientific or engineering argument without losing the plot. The app is where a person still sits in the loop on purpose. Hard analysis. Hard math. Hard trade studies. The jobs where you want the model to think, not wander off into a browser tab.
In Grok Bot, it should make the teammate framing less of a slogan. Bot already has a persistent computer, browser, filesystem, terminal, and plugins into Outlook, Calendar, and OneDrive. A father used it to help ship an Apple Watch app for newborn bath temperature, including navigating App Store submission in the browser. That is the product in one example: a real problem, a finished artifact, and an agent moving through human interfaces instead of leaving the user stuck in a chat loop.
The limit on Bot is not only the interface. It is whether the model can stay on a job, use the tools correctly, and come back only when a decision is required. A more token-efficient 4.7 is the difference between an agent that dies mid-task and an agent that finishes the inbox, the brief, or the form. Inbox zero only counts if the important five still reach a human. 4.7 does not remove that rule. It should make the sorting cheaper.
In Grok Build, it should push the coding loop from good autocomplete to close the ticket. 4.6 already took CursorBench. People have used Build to boot a game, record it, edit the footage, and voice a trailer without touching the pipeline. 4.7, if it is better in every way that is not latency, should reduce the number of times a human has to rescue a half-finished repo. The useful test is not a generated landing page. It is a repository with tests, an error that is not in the first file you open, and a model that can keep state across the repair.
In Grok Voice, it should tighten the speech-to-action loop. Voice Think Fast 2.0 already ranked at the top of speech-to-speech agent tests and is handling Starlink support volume at scale: tens of thousands of inbound calls a day, hardware diagnosis, replacements, thousands of orders a week. Elon’s own comment on the product was short. Grok Voice 2 is great. A stronger reasoner behind the voice is how a call stops being a transcript and becomes a completed order.
None of that requires science fiction. It requires the same products, a larger pretrain, and a corpus that knows machines.
The physics angle is the one that will decide whether the engineering claim is real.
A model that can hold GPQA Diamond is a model that can keep units, rates, and mechanisms in working memory. Orbital compute is the next public test of that muscle. SpaceX and Nvidia have already described a space-optimized Vera Rubin NVL72 path toward orbit, with a first generation of Starmind infrastructure aimed at taking the same accelerated computing architecture off the ground. The naive objection is always cooling. In vacuum you do not blow air over a rack. You radiate heat. The serious version of the objection names radiator rejection per square meter, coolant temperature, and the maximum operating temperature of the chip.
Elon’s reply to the expert-opinion version of that debate was that successful hardware in the field is worth an infinite number of expert opinions, and that people calling orbital AI impossible often do not know the basic thermal numbers. A 4.7 trained into real engineering should be able to stay inside that argument. Not by cheering. By keeping the units straight.
Starship is the same class of problem. Stainless steel was chosen because it gets stronger at cryogenic temperature, takes reentry heat better than the fashionable composite, welds at industrial speed, and does not need paint. Carbon fiber looked advanced and was slow, expensive, and poorly matched to a vehicle that must survive both liquid oxygen cold and reentry heat. A 4.7 that is actually better at real-world engineering should be able to reason about trades like that without collapsing into slogan.
The second pad at Starbase, the Louisiana site planned for high cadence, Falcon Heavy contracts running through about 2030 with the hope that payloads move to Starship, a full-duration 33-engine static fire, Ship 40 inspected in the water with heat-shield samples in hand: these are not metaphor. They are the environment the model is being asked to understand. If supplemental SpaceX data means anything, it means the model has seen more of that environment than a model trained only on the open web.
FSD is the terrestrial version. Fourteen billion supervised miles is a dataset accumulating at a pace that adds a billion miles in weeks, not years. The model that understands kinematics, friction, and edge cases is the model that can talk about driving as physics instead of as vibes. Tesla’s software stack already puts Grok in the cabin for calls, climate, music, and vehicle questions. A stronger reasoner makes that cabin assistant less of a novelty and more of a layer that can answer from the same world the car is driving through.
Starlink is the communications layer under all of it. High-altitude research aircraft asking for a kit that holds high-speed connectivity above 50,000 feet. A robot landing in the 2030s that only works if a minimally viable Starlink and StarMind constellation is already there. Voice agents resolving support at scale. The network is not a side quest. It is how the later machines keep talking after they leave the factory.
Predictions, labeled as predictions.
4.7 will be compared first on GPQA Diamond, CursorBench, the Agentic Index, and healthcare agent benches, because those are the boards Elon has already amplified. If it is a true upgrade, it should not lose first on scientific reasoning and should move the coding and agent scores without a six-times cost explosion.
4.7 will feel slightly slower in interactive chat and stronger on jobs that run for minutes instead of seconds. That is the trade Musk already described. People who judge a model by the first token will call it worse. People who judge a model by whether the job finished will call it better.
4.7 will be most obviously better on engineering prompts that involve hardware, operations, and constraint satisfaction. If the SpaceX data did anything, it will show there. Ask for a thermal balance, a weld sequence, a failure tree, a test that would kill a bad design. Keep the same prompt across models.
4.7 will not erase 4.6. Fast and cheap models stay in the stack. The new model takes the hard lane. That is how 4.5 and 4.6 already coexist with Voice and Bot variants. The family gets denser, not narrower.
4.7 will be judged less by a launch video and more by whether Bots finish work, Builds compile, and Voice calls resolve. The father with the Watch app is the template. The support call that ships a replacement is the template. The coding agent that does not need a second human to unstick it is the template.
4.7 will not make work optional by itself. Musk’s longer claim is that AI and robots will be able to do everything, that work will be optional, and that universal high income follows from output rising faster than claims on that output. 4.7 is one step on that curve, not the end of the curve. The remaining scarce skill is still judgment: what to assign, what not to permit, what to verify before money moves or code ships.
What to verify when it actually ships.
The model ID and the public card. Context window. Pricing. Whether “slightly slower” is ten percent or two times. Independent GPQA Diamond, CursorBench, and agent scores, not only vendor charts. Whether Bot and Build get the new weights on day one or trickle. Whether Voice stays on a fast variant while the heavy reasoner sits behind the hard tasks. Whether the engineering advantage is visible to people who do not work at SpaceX.
Until those numbers exist, 4.7 is a scheduled object with a stated size, a stated trade, and a stated corpus. That is already more signal than most unreleased models get. It is still not a finished scoreboard.
The deeper point is the cadence.
4.5, then 4.6, then 4.7. The company is not waiting a year to swing at the frontier. It is treating the model like Starship hardware: fly, measure, train, fly again. 4.6 proved the post-training recipe could take first on science and coding efficiency at a price that undercuts several rivals. 4.7 is the larger vehicle.
If the ten-day clock holds, the next week is not for mythology. It is for preparing the tests. Put the same engineering question to 4.6 today. Put it to 4.7 the morning it lands. Keep the prompt. Keep the constraints. Score the result the way you would score a static fire: did it run the full duration, or did it shut down early.
A good test is not “write me a poem about Mars.” A good test is a closed physical system with numbers. Radiator area. Heat load. Chip temperature limit. A tank at cryogenic temperature. A procedure with a forbidden step. A codebase with a failing test and a dependency that is not in the README. A support transcript that requires a tool call, not a sympathetic paragraph.
That is how you find out whether 2.1 trillion parameters and a SpaceX corpus were a press line or a better reasoner.
Grok 4.6 already showed that a reasoner can lead scientific and agent boards while remaining cheap enough to run as daily infrastructure. Grok 4.7 is the attempt to keep that lead after the next wave of rival models, with more capacity and a corpus pulled from machines that actually fly.
The stack is already built. The clock is public. The only honest review will be hardware in the field: a model that finishes the job.
Published on grokbot.sh. Cite the public log, not a prompt pack.