InsightsPublished 8 min read

Gemini 4 Argon: Google's new frontier model, what it scores and when you can use it

Google's first new flagship in months matches OpenAI's best on independent testing, leads its own table on legal, finance and business-workflow benchmarks, and is far more willing to say it does not know. The catch: almost nobody can use it yet.

By Monolith

Three security analysts at night looking together at a wall of monitors showing abstract blue data patterns
A small security operations team reviewing activity on a wall of monitors at night.

Imagine a mid-size accounting firm that has been waiting for AI good enough to trust with real client work: reading a stack of financial statements, drafting a research memo, checking a contract clause against the firm's standard terms. The two things that have held it back are quality on specialist tasks and answers that sound right but are made up. Google's new model is aimed at both, and on the second one the independent numbers are striking.

Google announced Gemini 4 Argon on September 30, 2026, calling it its next era of frontier intelligence. It is built for long, multi-step work in software engineering, enterprise knowledge work such as legal and finance, and cybersecurity defense. But it is not generally available. For now it is rolling out only to a set of trusted cyber defenders, with developers, businesses and consumers to follow once Google finishes strengthening its safeguards. 1

Who can use Gemini 4 Argon today

Almost nobody outside Google. Argon is rolling out first through Google's Fairwind Program, which gives high-priority defenders such as governments, healthcare providers and telecommunications services early access to advanced models for cybersecurity work. Fairwind has more than 650 partners, but only a set of them get Argon, and organizations may grant it only to internal cybersecurity, incident response or penetration testing teams. 14

Google says it is taking part in the U.S. government's voluntary pre-release access process and will gather feedback from early testers before opening Argon up, starting with paid API customers and Google AI Ultra subscribers. It has not given a date. For a business, that means Argon is something to plan for, not something to buy this week. 1

Access sequence
  1. Sept 30, 2026Trusted cyber defenders (Fairwind)Rolling outSelected Fairwind partners, for internal security, incident response and penetration testing teams only.
  2. NextPaid API customers and Google AI UltraNo date yetThe first wider release, after Google strengthens its safeguards with early-tester feedback.
  3. LaterDevelopers, enterprises and consumersNo date yetBroad availability, which Google says will come as soon as possible.
Rollout order as described in Google's announcement and Fairwind Program page, September 30, 2026. No dates have been given beyond the first step. [1][4]

What Google's own benchmarks show

Google published a table of 18 benchmarks comparing it with GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5. Argon posts the top score on most of them, and the biggest margins are in the areas most relevant to professional services: Harvey's Legal Agent Benchmark, Vals Finance Agent v2 and AutomationBench, Zapier's test of end-to-end business workflows. Argon also sets a new best of 77.9% on DeepSWE v1.1, a test of long real-world software engineering tasks. 1

The table is honest about where Argon does not lead, and those rows matter just as much. GPT-6 Astra is ahead on FrontierSWE, a harder coding test, on OSWorld 2.0 for operating a computer, and on scientific research workflows. Claude Opus 5.5 leads Terminal-Bench 4.0 and machine-learning engineering. On CWE-bench, a test of fixing security vulnerabilities, Argon and Astra tie at 68%. 1

Read the vendor table with one caution. Google's methodology page says Argon was run through the Gemini API at its highest thinking setting, several Argon scores were computed by Google itself, and the other models' numbers come from their makers' own reports. That is the normal way launch tables are built, but it is not a controlled head-to-head. The independent results below are the better guide. 2

What independent testing says

Artificial Analysis runs every model through the same ten evaluations. On its Intelligence Index, Gemini 4 Argon at high reasoning scores 53, level with GPT-6 Astra at maximum effort and one point ahead of GPT-6.1 Sol. It is 23 points above Google's previous non-Flash model, Gemini 3.1 Pro Preview. Claude Opus 5.5 and Claude Sonnet 5.5 still sit higher, at 58 and 56. 3

Cost is where the launch discount does the work. At the introductory price, Artificial Analysis puts Argon at $1.99 per index task, about 60% of GPT-6 Astra's $3.26, but nearly three times GPT-6.1 Sol's $0.72. The saving comes from lower token prices, not efficiency: Argon averaged about 62,000 output tokens per task, more than twice Astra's 27,000. When Google's price doubles after the launch period, the same work would cost about $3.98 by simple doubling, slightly more than Astra. 13

ModelInputCached inputOutputCost per AA task
Gemini 4 Argon (launch)$2.00$0.10$10.00$1.99
Gemini 4 Argon (standard)$4.0095% off input$20.00about $3.98
GPT-6 Astra$10.00$1.00$50.00$3.26
GPT-6.1 Sol$2.00$0.10$10.00$0.72
Claude Opus 5.5$4.00$0.20$20.00$5.98
API prices in USD per million tokens. Gemini 4 Argon from Google's announcement; rival list prices as shown by Artificial Analysis, read September 30, 2026. Google has not said when the introductory period ends. 13

The standout: far fewer made-up answers

The most useful number for a business may be this one. On Artificial Analysis's AA-Omniscience test, which asks hard knowledge questions and penalizes confident wrong answers, Argon's hallucination rate is 15%, far below the other leading models: GPT-6 Astra's is 51% and GPT-6.1 Sol's 54%. In practice, when Argon does not know something, it is much more likely to say so than to guess. 3

There is a trade-off. Argon also answers fewer questions correctly: its accuracy on the same test is 50%, against 63% for GPT-6 Astra. Its overall AA-Omniscience score, which balances the two, lands at 42, level with Astra's 43 and GPT-6.1 Sol's 42. So Argon does not know more; it is better at knowing what it does not know. For legal, finance and client-facing work, where a confident wrong answer is expensive, that is often the better trade. 3

Two professionals at an office table reviewing printed contracts beside a laptop, one pointing to a highlighted paragraph
Two professionals checking a contract clause together in a small law and accounting office.

Stronger at getting work done

Agentic work, where a model carries out a multi-step task with tools, has historically been a weaker area for Gemini. Artificial Analysis's results show that changing. Argon ranks first on its version of AutomationBench at 77.5%, ahead of Claude Sonnet 5.5 at 71.3%, and scores 57% on Terminal-Bench 4.0, up from 4% for Gemini 3.1 Pro Preview. 3

Google's own examples point the same way. Inside Google, teams of Argon agents analyzed fleet-wide data to find memory optimizations that freed more than 300 TiB of memory, and Argon agents are migrating large C and C++ codebases to the Rust language, with human review before anything reaches production. Google also raised the output limit to 1 million tokens per response, so the model can work through very long reasoning or generate large documents in one go. 1

Security first, by design

Google trained Argon to be highly capable at cybersecurity defense: finding, validating and patching software vulnerabilities on its own. For trusted defenders and Google's internal teams, it will be released without its cyber guardrails. That is exactly why access is limited. Google describes four safeguard areas it is strengthening first: defending against misuse, resisting prompt injection attacks, monitoring the model's reasoning for signs it is overstepping the user's intent, and hardening the test environments. 1

Prompt injection is the one to understand if you plan to connect AI to email, documents or the web. It is when hidden instructions in content the model reads try to hijack what it does. Google says Argon is its most resilient model yet against these attacks. That lowers the risk; it does not remove the need for approval steps on anything that sends, pays or deletes. 1

What this means for your business

If you are...What Argon changesWhat to do now
A law, accounting or finance firmLeading vendor scores on legal and finance agents, and low hallucinationList two research or review tasks to trial when access opens
Automating business workflowsTop AutomationBench results on both Google's and AA's versionsMap one workflow now so it is ready to test
Mostly writing codeStrong on DeepSWE; behind on FrontierSWE and Terminal-BenchKeep your current coding model; compare on your own repo later
Cost-sensitive at high volumeLaunch price is attractive but doubles later; Sol is cheaper per taskPrice on the standard rate, not the discount
A security teamDesigned for defense; Fairwind access for qualifying organizationsCheck whether your organization qualifies for Fairwind
Monolith's suggested planning moves, based on Google's announcement and Artificial Analysis's results. Suggestions, not results of Monolith testing. 13

Two cautions for planning. First, the launch discount is temporary and Google has not said when it ends, so any budget should assume the standard $4 and $20 prices. Second, Argon uses a lot of output tokens per task, which the per-token price hides. Test on your own work and compare total cost per finished task, not the price per million tokens. 13

Comparing it with OpenAI's newest mid-priced model? Our DevDay guide covers GPT-6.1 Sol's price and results.

The short version

Monolith's take: Gemini 4 Argon puts Google back among the top AI labs, with leading results on legal, finance and business-workflow tests and a hallucination rate far below its rivals. It is not available to most businesses yet, and its price doubles after launch. Prepare the workflows you would trust it with, and judge it on cost per finished task when access opens.

Questions, answered
When can I use Gemini 4 Argon?
Not yet for most people. It is rolling out to selected cyber defenders in Google's Fairwind Program first, then to paid API customers and Google AI Ultra subscribers, with no date announced.
How much does Gemini 4 Argon cost?
Google's introductory price is $2 per million input tokens and $10 per million output tokens, with cached input 95% off. After the introductory period it rises to $4 and $20.
Is Gemini 4 Argon better than GPT-6 Astra?
They tie at 53 on Artificial Analysis's Intelligence Index. Argon leads on legal, finance and workflow tests and hallucinates far less; Astra leads on some coding, computer-use and science tests.
Sources

Read for this feature. The numbers match the markers in the text.

  1. Google: Gemini 4 Argon, our next era of frontier intelligenceblog.google
  2. Google DeepMind: Gemini 4 Argon model evaluation methodologydeepmind.google
  3. Artificial Analysis: Gemini 4 Argonartificialanalysis.ai
  4. Google DeepMind: Fairwind Programdeepmind.google

Planning which AI models to trust with client work, and on which tasks? We help teams test and choose.

Where this leads

Want to apply this to a project? Here is the related service.

Free download

25 AI instructions to try on everyday work

Prompts are instructions you give an AI tool. This PDF includes 25 examples for inquiries, marketing and admin. Adapt them to your task and check the results.

Enter your email to access the PDF. This form does not sign you up for a marketing sequence.