InsightsUpdated 5 min read

Grok 4.6 and Grok Bot: xAI's play for the always-on coworker

In one week xAI shipped a frontier-class model and a product that signs into your software and works while you sleep. One of those is a pricing story. The other is a permissions story.

A business operator reviews an AI task queue and approval panel on a wide monitor
Always-on only works when a human still owns the approval

xAI released Grok 4.6 on August 12, 2026, one day after Grok Bot entered public beta. Taken together they are the clearest statement yet of where xAI wants to sell: not a chat window, a workforce. The model is a genuine step up from 4.5. The bot product is the more interesting and the more dangerous of the two.

Grok 4.6: the numbers, with their sources

On the Artificial Analysis Intelligence Index, an independent aggregate, Grok 4.6 scores 61, tied with GPT-5.6 Sol and one point behind Claude Fable 5. The jumps over its own predecessor are large: DeepSWE from 54 to 65.9, APEX-Agents from 47.1 to 57.5. It also tops the GDPVal-AA row, an economic-value benchmark, at 1753. Context stays at 500,000 tokens, there are now four reasoning-effort levels with a new xhigh, and the knowledge cutoff is February 1, 2026.

BenchmarkGrok 4.6GPT-5.6 SolClaude Fable 5
AA Intelligence Index616162
DeepSWE v1.165.9%73%not listed
APEX-Agents57.5%not listed59.2%
CursorBench v3.269.9%not listed70.5%
FrontierCode v1.161.3%not listed63.6%
Frontier comparison on independent benchmarks, as compiled by GEO Toolbox from Artificial Analysis data at launch. Higher is better.

The honest read from that table is the one GEO Toolbox gave: Fable wins more rows than anyone, most of Grok's wins are inside the noise, and Grok 4.6 is a credible mid-priced frontier model rather than the top of the index. The word doing the work there is mid-priced.

Update, September 3, 2026: Claude Fable 5.1 shipped on September 1 and is not in this table. The only shared row published so far is Terminal-Bench 4.0, where Anthropic lists 5.1 at 55.8 and GPT-5.6 Sol at 37.3, as reported by Tech-ish; xAI has not published a Grok 4.6 figure on it. Read the Fable column above as the June model.

Two dollars in and six out puts Grok 4.6 level with GPT-5.6 Luna on output price while scoring like Sol. That is the pitch, and for prompts under 200,000 tokens it is a real one. Cross that line and the entire request bills at double, which erases the advantage for anyone stuffing a whole knowledge base into every call. Availability at launch was API, Grok Build, Cursor, Microsoft Office add-ins, and the usual gateways; the consumer chatbot was not on the launch list.

Grok Bot: an agent with its own computer

Grok Bot launched in beta on August 11. The idea is simple to state and hard to do well: you create a persistent Bot with a job, give it access to the applications and websites it needs, and it works through multi-step tasks on its own computer in the cloud, with a browser, a filesystem, and a terminal. It signs into your tools with your credentials, keeps going when your laptop is closed, and comes back when it needs approval or has finished. Bots coordinate in shared threads, learn routines from demonstrations, and carry memory across conversations.

PlanPriceWho it is for
Cursor Ultra$200Individual builders already living in Cursor
Cursor Teams Premium$120 per seatTeams that want Bots alongside shared coding agents
SuperGrok Heavy$300xAI's own top consumer and prosumer tier
How to get Grok Bot, per VentureBeat's launch coverage. All prices monthly.

The part worth pausing on: Bots work through the user interface, which means they can operate legacy software and internal systems that have no API and never will. For a lot of local businesses, that is the entire back office. It also means a Bot with the wrong permissions can edit CRM records, answer support tickets, and change vendor orders exactly as confidently as it does the right things. VentureBeat flagged the operational risk plainly, and so will anyone who has watched an automation misfire against production data.

SystemReadDraftExecuteNever
CRMYesNotes and follow-up draftsStage changes, with approvalDelete records
Support deskYesReply draftsSend, with approvalRefunds
Vendor portalYesOrder draftsNonePlace orders
AccountingSummaries onlyNoneNonePayments
Monolith's starting permission matrix for an always-on Bot. Every Execute cell carries a human approval; the Never column is not negotiable in the first ninety days.

The approval-gates field guide covers the risk ladder behind that matrix, from read-only to irreversible, and the four tests a gate has to pass.

What this means for agencies and businesses

  • Model choice is now a cost lever, not a loyalty. Grok 4.6 under 200K tokens is a legitimate default for research, drafting, and agentic runs that do not need the last point on the index.
  • Always-on agents need an approval design before they need a subscription. Which actions run unattended, which pause for a human, and what a Bot is never allowed to touch is the whole project.
  • The UI-driving approach is the unlock for shops running old software. It is also the reason to scope permissions per Bot, per system, and test against a copy first.

Monolith's AI strategy and consulting engagement produces exactly that document: which model runs which workload, which workflows get an always-on agent, where the approval gates sit, and what it should cost. Quoted by scope, vendor-neutral.

Not sure which workflow to hand to an agent first? The 60-second automation check scores the candidates, no email required.

Questions, answered
How much does Grok 4.6 cost?
$2 per million input tokens and $6 per million output for requests under 200,000 tokens; $4 and $12 above that line, applied to the whole request rather than the overage. Context is 500,000 tokens.
What is Grok Bot?
xAI's persistent agent product, in public beta since August 11, 2026. A Bot has a job, signs into your applications with your credentials, and works through multi-step tasks on its own cloud computer with a browser, a filesystem, and a terminal, returning when it needs approval or has finished.
Is Grok Bot safe to connect to business systems?
Only as safe as its permissions. Because it drives the user interface, it can change CRM records or vendor orders as confidently as it reads them. Scope read, draft, and execute per system, keep money and deletions off the table, and test against a copy first.
The short version

Grok 4.6 is frontier scores at a mid-tier price, with a cliff at 200K tokens. Grok Bot is the always-on coworker every vendor promised, and its value is set entirely by how carefully you scope what it is allowed to do.

Where this leads

This is the thinking behind a service Monolith runs every week.

Free download

The prompt library every small team should have

25 copy-paste prompts for leads, marketing, operations, hiring, and admin. The ones we actually give clients, as a designed PDF. Free, in exchange for an email.

Instant download, no drip sequence. Unsubscribe is not needed because nothing else gets sent.