Grok 4.6 and Grok Bot: xAI's play for the always-on coworker
In one week xAI shipped a frontier-class model and a product that signs into your software and works while you sleep. One of those is a pricing story. The other is a permissions story.

xAI released Grok 4.6 on August 12, 2026, one day after Grok Bot entered public beta. Taken together they are the clearest statement yet of where xAI wants to sell: not a chat window, a workforce. The model is a genuine step up from 4.5. The bot product is the more interesting and the more dangerous of the two.
Grok 4.6: the numbers, with their sources
On the Artificial Analysis Intelligence Index, an independent aggregate, Grok 4.6 scores 61, tied with GPT-5.6 Sol and one point behind Claude Fable 5. The jumps over its own predecessor are large: DeepSWE from 54 to 65.9, APEX-Agents from 47.1 to 57.5. It also tops the GDPVal-AA row, an economic-value benchmark, at 1753. Context stays at 500,000 tokens, there are now four reasoning-effort levels with a new xhigh, and the knowledge cutoff is February 1, 2026.
| Benchmark | Grok 4.6 | GPT-5.6 Sol | Claude Fable 5 |
|---|---|---|---|
| AA Intelligence Index | 61 | 61 | 62 |
| DeepSWE v1.1 | 65.9% | 73% | not listed |
| APEX-Agents | 57.5% | not listed | 59.2% |
| CursorBench v3.2 | 69.9% | not listed | 70.5% |
| FrontierCode v1.1 | 61.3% | not listed | 63.6% |
The honest read from that table is the one GEO Toolbox gave: Fable wins more rows than anyone, most of Grok's wins are inside the noise, and Grok 4.6 is a credible mid-priced frontier model rather than the top of the index. The word doing the work there is mid-priced.
Update, September 3, 2026: Claude Fable 5.1 shipped on September 1 and is not in this table. The only shared row published so far is Terminal-Bench 4.0, where Anthropic lists 5.1 at 55.8 and GPT-5.6 Sol at 37.3, as reported by Tech-ish; xAI has not published a Grok 4.6 figure on it. Read the Fable column above as the June model.
Two dollars in and six out puts Grok 4.6 level with GPT-5.6 Luna on output price while scoring like Sol. That is the pitch, and for prompts under 200,000 tokens it is a real one. Cross that line and the entire request bills at double, which erases the advantage for anyone stuffing a whole knowledge base into every call. Availability at launch was API, Grok Build, Cursor, Microsoft Office add-ins, and the usual gateways; the consumer chatbot was not on the launch list.
Grok Bot: an agent with its own computer
Grok Bot launched in beta on August 11. The idea is simple to state and hard to do well: you create a persistent Bot with a job, give it access to the applications and websites it needs, and it works through multi-step tasks on its own computer in the cloud, with a browser, a filesystem, and a terminal. It signs into your tools with your credentials, keeps going when your laptop is closed, and comes back when it needs approval or has finished. Bots coordinate in shared threads, learn routines from demonstrations, and carry memory across conversations.
| Plan | Price | Who it is for |
|---|---|---|
| Cursor Ultra | $200 | Individual builders already living in Cursor |
| Cursor Teams Premium | $120 per seat | Teams that want Bots alongside shared coding agents |
| SuperGrok Heavy | $300 | xAI's own top consumer and prosumer tier |
The part worth pausing on: Bots work through the user interface, which means they can operate legacy software and internal systems that have no API and never will. For a lot of local businesses, that is the entire back office. It also means a Bot with the wrong permissions can edit CRM records, answer support tickets, and change vendor orders exactly as confidently as it does the right things. VentureBeat flagged the operational risk plainly, and so will anyone who has watched an automation misfire against production data.
| System | Read | Draft | Execute | Never |
|---|---|---|---|---|
| CRM | Yes | Notes and follow-up drafts | Stage changes, with approval | Delete records |
| Support desk | Yes | Reply drafts | Send, with approval | Refunds |
| Vendor portal | Yes | Order drafts | None | Place orders |
| Accounting | Summaries only | None | None | Payments |
The approval-gates field guide covers the risk ladder behind that matrix, from read-only to irreversible, and the four tests a gate has to pass.
What this means for agencies and businesses
- Model choice is now a cost lever, not a loyalty. Grok 4.6 under 200K tokens is a legitimate default for research, drafting, and agentic runs that do not need the last point on the index.
- Always-on agents need an approval design before they need a subscription. Which actions run unattended, which pause for a human, and what a Bot is never allowed to touch is the whole project.
- The UI-driving approach is the unlock for shops running old software. It is also the reason to scope permissions per Bot, per system, and test against a copy first.
Monolith's AI strategy and consulting engagement produces exactly that document: which model runs which workload, which workflows get an always-on agent, where the approval gates sit, and what it should cost. Quoted by scope, vendor-neutral.
Not sure which workflow to hand to an agent first? The 60-second automation check scores the candidates, no email required.
How much does Grok 4.6 cost?
What is Grok Bot?
Is Grok Bot safe to connect to business systems?
Grok 4.6 is frontier scores at a mid-tier price, with a cliff at 200K tokens. Grok Bot is the always-on coworker every vendor promised, and its value is set entirely by how carefully you scope what it is allowed to do.
This is the thinking behind a service Monolith runs every week.
The prompt library every small team should have
25 copy-paste prompts for leads, marketing, operations, hiring, and admin. The ones we actually give clients, as a designed PDF. Free, in exchange for an email.