Grok 4.7: from a good answer to a useful finished task
A more capable model becomes valuable when it helps finish something you can actually use. Here is how to read this release, choose a sensible first task and judge the result.
By Monolith

Picture a small coffee company preparing a tasting event. Someone needs to turn a rough brief into a booking page, check the details and make sure the form works on a phone. That job contains writing, design decisions and technical checks. It is also a useful way to understand what people mean when they say AI is getting better at doing longer tasks.
Grok 4.7 launched on September 21, 2026. SpaceXAI, the company behind Grok, says it is built on a new, larger base model and was given a longer reinforcement learning run, the training stage where a model practices tasks and is rewarded for good results, weighted toward problems that take many hours to complete. The company says the result works longer on difficult tasks and checks its own work more carefully. The release benchmarks are evidence about particular tests; they do not establish how well it will handle your event page or client account. 1
First, what is an AI agent?
A normal chat exchange might end with instructions for building the page. An agent is an AI assistant connected to tools that let it take steps toward completing the work. With suitable access, those tools might let it edit files, open a preview or run a test. The surrounding software determines what it can actually do. A capable model inside an ordinary chat box does not automatically have access to your website.
Think of the model as the part deciding what to try next. The workspace supplies the materials and tools. Your brief supplies the purpose and boundaries. If any of those are missing, a strong model can still produce a polished answer that solves the wrong problem.
What changed, and what do the scores mean?
Two published coding results help illustrate the change from Grok 4.6. A benchmark is a repeatable set of tasks used to compare systems. The percentages below describe success on those tests, rather than a prediction about the percentage of your work an assistant will finish. The release compares different reasoning settings: 4.7 at xhigh and 4.6 at high. This is the publisher's comparison, not Monolith's independent test. 1
For an agency, the useful next question is concrete: can the new model make the same requested change with fewer corrections? Try both versions on a copy of one small project. Count broken requirements, time spent reviewing and whether the result passes your checks. A higher public score is a reason to run that comparison; your own finished work is the deciding evidence.
How it compares with GPT and Claude, on SpaceXAI's own chart
The announcement also lines Grok 4.7 up against two rivals: OpenAI's GPT-5.6 Sol and Anthropic's Claude Fable 5.1, each at its highest reasoning setting. The table is worth reading row by row, because the picture is mixed rather than a clean win. 1
| Measure | Grok 4.7 | GPT-5.6 Sol | Claude Fable 5.1 |
|---|---|---|---|
| Price per 1M tokens, in / out | $2 / $6 | $4 / $20 | $10 / $50 |
| CursorBench 4.0 | 46.3% | 41.7% | 51.8% |
| DeepSWE v1.1 | 71.0%* | 72.7% | 70.0% |
| Terminal-Bench 4.0 | 37.6% | 37.3% | 57.9% |
| EEBench | 64.0% | 39.4% | 56.4% |
| Harvey Legal Agent | 19.6% | 2.5% | 6.7% |
| HealthBench Pro | 56.7% | 60.5% | 62.1% |
| AA-Briefcase (Elo) | 1,657 | 1,487 | 1,678 |
Three things stand out. Grok 4.7 leads on the engineering and legal tests, at a fraction of the per-token price. Claude Fable 5.1 still leads on the hardest coding and terminal work, and on clinical reasoning. And a lead is not the same as reliability: 19.6% on the legal benchmark means roughly four in five of those legal agent tasks did not pass, so a lawyer still does the work that matters. Notice also who is missing. The publisher chose the comparison, and two of the strongest current models, Claude Opus 5 and GPT-6 Astra, are absent from the main table. 1
What independent testing found
Artificial Analysis is an independent firm that runs the same tests across many models, which makes it a useful check on a publisher's numbers. It evaluated Grok 4.7 on launch day and publishes its scores next to the major rivals. Some of its tests report an Elo score, a head-to-head rating borrowed from chess: a higher number means the model's work tends to win when compared directly with other models. 45
| Measure | Grok 4.7 | Grok 4.6 | Claude Fable 5.1 | Claude Opus 5 | GPT-6 Astra |
|---|---|---|---|---|---|
| Intelligence Index | 46 | 44 | 53 | 51 | 53 |
| GDPval-AA (Elo) | 1,695 | 1,605 | 1,735 | 1,708 | 1,542 |
| AA-Briefcase (Elo) | 1,657 | 1,546 | 1,678 | 1,673 | 1,569 |
| Speed (tokens/s) | 39 | 59 | 66 | 55 | 61 |
On the broad Intelligence Index, Grok 4.7 gained two points over its predecessor and remains behind the leading Anthropic and OpenAI models. Its real progress is in agentic knowledge work, the long, multi-step office tasks this article is about: it gained 111 Elo on AA-Briefcase and 90 on GDPval-AA, which puts it close behind Claude Opus 5 and Claude Fable 5.1 on both. On the firm's coding agent measure, Grok 4.7 running in Grok Build ranked fourth, behind Fable 5.1, GPT-6 Astra and Opus 5. 467
The detail tells you where not to over-read. Outside knowledge work, Artificial Analysis found Grok 4.7 broadly matched Grok 4.6. On Terminal-Bench 4.0 it measured a gain of 4.5 points, far smaller than the 17.3 points in SpaceXAI's own results for the same pair of settings. The two used different test setups, which is exactly why an independent run is worth checking. It also recorded small declines on its long-context reasoning test and on AutomationBench. A 500,000-token window tells you how much material fits; it does not guarantee better reasoning across all of it. 4
You may also have seen posts on X saying Grok 4.7 beat GPT-6 Astra on professional work. On these two work leaderboards, that is accurate. It is also partial: on the same leaderboards, Claude Fable 5.1 and Claude Opus 5 both score higher. A comparison that names a single rival is always worth checking against the full table. 67
Why the same price can mean a bigger bill
Grok 4.7 keeps Grok 4.6's per-token prices. Cost per finished task is a different number. Artificial Analysis found that at xhigh effort Grok 4.7 generated about 81,000 output tokens per Intelligence Index task, against 36,000 for Grok 4.6 at high and 27,000 for GPT-6 Astra at max: 125% and 196% more. Output tokens include the model's working-out, and they are billed at the output rate. A cheaper price per token can still produce a larger bill per result. 34
SpaceXAI's own chart argues the opposite for coding. Plotting CursorBench score against average cost per task, it places the Grok 4.7 curve above Claude Opus 5, GPT-5.6 Sol and Claude Sonnet 5 at the same spend, with Fable 5.1 reaching the highest score at a higher cost. Both findings can be true. They measure different tasks, in different tools, at different settings. Only a test on your own work tells you which one applies to you. 1
Speed deserves the same care. The announcement headline promises twice the speed at half the price of comparable models, without naming the models or settings. Artificial Analysis measured Grok 4.7 at xhigh producing about 39 tokens per second, slower than Grok 4.6 and each rival in the table above. The separate Fast variant doubles output speed at double the price. 15
Try high effort before xhigh
Grok 4.7 lets you choose how hard it reasons: low, medium, high or xhigh, with high as the API default. The headline results use xhigh. Artificial Analysis tested both, and on the knowledge work that matters most here, the difference was small. 2567
| Setting | Intelligence Index | GDPval-AA (Elo) | AA-Briefcase (Elo) | Speed (tokens/s) |
|---|---|---|---|---|
| xhigh | 46 | 1,695 | 1,657 | 39 |
| high | 46 | 1,694 | 1,644 | 47 |
On both work benchmarks the published margins of error for the two settings overlap, so the results cannot separate them, while high ran about a fifth faster. For most business tasks, start at high and move up only when a specific job clearly benefits. Less reasoning usually means fewer billed tokens as well, though the token figures above are for xhigh, so measure your own. 67
Four jobs worth trying first
| Job | Give it | Ask for | Check yourself |
|---|---|---|---|
| Campaign page | Approved offer, brand files and a safe project copy | A preview and a list of completed checks | Mobile layout, links, form delivery and offer accuracy |
| Customer research | Approved questions and access to public sources | A short brief with dated source links | Whether the sources actually support each conclusion |
| Document or slide deck | An outline, approved facts and a template | A first draft with a source for every figure | Each number and claim against the source material |
| Reporting helper | A small, cleaned export and metric definitions | A repeatable report plus its calculations | Totals, missing rows and the date range |
| Internal tool | One repetitive task and sample inputs | A working prototype with a simple handover | Permissions, error cases and who will maintain it |
Keep the first job narrow enough to inspect. Building one booking page teaches you more than asking an assistant to improve the entire business. You will discover whether the bottleneck is the model, incomplete information, unclear instructions or a process nobody has properly defined yet.
The release also hints at where to test first. SpaceXAI says Grok 4.7 is better at creating documents and presentations, and its strongest relative scores were in electrical engineering and legal agent work. Treat those as reasons to try it on a document-heavy or specialist drafting task with an expert reviewing the output, rather than as proof it will do that work unsupervised. 1

A brief that gives the assistant a fair chance
For our imaginary coffee event, try this: Create a draft booking page using the attached event details and brand materials. Explain the event in everyday language. Include the date, location, price and booking button. Ask me about missing details before inventing them. Check the page on a narrow screen, test the form in the preview environment, and list anything you could not verify. Stop before publishing.
That final handover matters. A screenshot can look finished while a form sends nowhere. Ask for evidence tied to the deliverable: which links were opened, which fields were tested and which details came from the supplied brief. Review time becomes much more useful when the assistant shows where uncertainty remains.
Which Grok 4.7 are you getting?
The standard model is available through the xAI API, through model gateways including OpenRouter, Vercel and Cloudflare, and in products including Cursor and Grok Build, where it is the default model for the coding agent. An API is a connection developers use to put a model inside other software. Grok 4.7 Fast uses the same model on faster infrastructure at twice the standard rate; current documentation limits it to Cursor and Grok Build, excluding Grok Build's free tier and the public API. Check the product you plan to use. 2
If your team already works in Cursor, the AI code editor, note how it treats Grok. Cursor lists Grok 4.7 among its own models, which draw on a separate usage pool with significantly more included usage than third-party models. Check Cursor's own pricing page for the long-context and Fast versions before you estimate a budget.
The model accepts text and images and produces text; making images needs a separate tool. Through the API it can call functions you define, search the web and X, and run code. Its documented knowledge cutoff is May 2026, so current research requires current material or connected search tools. Its 500,000-token context window describes how much material it can work with in one request. Tokens are the small pieces of text used to measure input and output; a large window is capacity, rather than a guarantee that every detail will be used correctly. 2
Safety settings and refusals
SpaceXAI says Grok 4.7 runs on an entirely new safeguard stack and is its strongest model yet at refusing harmful requests and resisting jailbreaks, the attempts to trick a model past its rules. In the company's own tests it topped LatchBio's biosafety benchmark at 62.4% and let through 3.3% of risky dual-use prompts on HackerBench, its cybersecurity test, while rarely blocking legitimate security work. It is also giving select security partners invite-only access to the model's red-team abilities for defensive research. These are company-reported results. 1
For a business, stricter defaults are usually welcome, but test your legitimate edge cases, such as a security review, a medical explainer or a legal summary, before building a workflow around the model. A refusal halfway through an automated job costs time too.
The cost includes what you send, what comes back and which tools run
| Prompt size | Input | Cached input | Output |
|---|---|---|---|
| Below 200,000 tokens | $2 | $0.50 | $6 |
| 200,000 tokens or more | $4 | $1 | $12 |
Input is material you send; output is material the model generates. Cached input is eligible reused material billed at a lower rate. Search can add another charge. Since noon Pacific on September 21, xAI's X Search has cost $5 per 1,000 posts fetched and $10 per 1,000 profiles fetched. A research task's cost therefore depends on what it retrieves, as well as the tokens it uses. These are API charges, not an all-inclusive subscription quote. 3
Monolith's take: Grok 4.7 is a real step up for long, multi-step office work at a per-token price well below its rivals, and it still trails the best Anthropic and OpenAI models on hard coding and broad reasoning. Start at high effort, test it on one real task, and measure the cost of an accepted result: the assistant's usage, your review time and the fixes needed before someone can use the work.
Is Grok 4.7 better than Claude or GPT?
Why do some reviews say Grok 4.7 costs more when the price did not change?
Can Grok 4.7 build my website?
Read for this feature. The numbers match the markers in the text.
- SpaceXAI: Grok 4.7 announcement, benchmark tables and safety results, September 21, 2026; checked September 22, 2026x.ai
- SpaceXAI: Grok 4.7 model, availability and Fast documentationdocs.x.ai
- SpaceXAI: API token and tool pricing; checked September 21, 2026docs.x.ai
- Artificial Analysis: Grok 4.7 evaluation results, posted on X, September 21, 2026x.com
- Artificial Analysis: model leaderboard, Intelligence Index and output speed; checked September 22, 2026artificialanalysis.ai
- Artificial Analysis: GDPval-AA leaderboard; checked September 22, 2026artificialanalysis.ai
- Artificial Analysis: AA-Briefcase leaderboard; checked September 22, 2026artificialanalysis.ai
Bring a real task to a workshop and learn how to brief, review and improve an AI-assisted workflow.
Want to apply this to a project? Here is the related service.
25 AI instructions to try on everyday work
Prompts are instructions you give an AI tool. This PDF includes 25 examples for inquiries, marketing and admin. Adapt them to your task and check the results.