Gemini 4 Argon: Google's new frontier model, what it scores and when you can use it
Google's first new flagship in months matches OpenAI's best on independent testing, leads its own table on legal, finance and business-workflow benchmarks, and is far more willing to say it does not know. The catch: almost nobody can use it yet.
By Monolith

Imagine a mid-size accounting firm that has been waiting for AI good enough to trust with real client work: reading a stack of financial statements, drafting a research memo, checking a contract clause against the firm's standard terms. The two things that have held it back are quality on specialist tasks and answers that sound right but are made up. Google's new model is aimed at both, and on the second one the independent numbers are striking.
Google announced Gemini 4 Argon on September 30, 2026, calling it its next era of frontier intelligence. It is built for long, multi-step work in software engineering, enterprise knowledge work such as legal and finance, and cybersecurity defense. But it is not generally available. For now it is rolling out only to a set of trusted cyber defenders, with developers, businesses and consumers to follow once Google finishes strengthening its safeguards. 1
Who can use Gemini 4 Argon today
Almost nobody outside Google. Argon is rolling out first through Google's Fairwind Program, which gives high-priority defenders such as governments, healthcare providers and telecommunications services early access to advanced models for cybersecurity work. Fairwind has more than 650 partners, but only a set of them get Argon, and organizations may grant it only to internal cybersecurity, incident response or penetration testing teams. 14
Google says it is taking part in the U.S. government's voluntary pre-release access process and will gather feedback from early testers before opening Argon up, starting with paid API customers and Google AI Ultra subscribers. It has not given a date. For a business, that means Argon is something to plan for, not something to buy this week. 1
- Sept 30, 2026Trusted cyber defenders (Fairwind)Rolling outSelected Fairwind partners, for internal security, incident response and penetration testing teams only.
- NextPaid API customers and Google AI UltraNo date yetThe first wider release, after Google strengthens its safeguards with early-tester feedback.
- LaterDevelopers, enterprises and consumersNo date yetBroad availability, which Google says will come as soon as possible.
What Google's own benchmarks show
Google published a table of 18 benchmarks comparing it with GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5. Argon posts the top score on most of them, and the biggest margins are in the areas most relevant to professional services: Harvey's Legal Agent Benchmark, Vals Finance Agent v2 and AutomationBench, Zapier's test of end-to-end business workflows. Argon also sets a new best of 77.9% on DeepSWE v1.1, a test of long real-world software engineering tasks. 1
The table is honest about where Argon does not lead, and those rows matter just as much. GPT-6 Astra is ahead on FrontierSWE, a harder coding test, on OSWorld 2.0 for operating a computer, and on scientific research workflows. Claude Opus 5.5 leads Terminal-Bench 4.0 and machine-learning engineering. On CWE-bench, a test of fixing security vulnerabilities, Argon and Astra tie at 68%. 1
Read the vendor table with one caution. Google's methodology page says Argon was run through the Gemini API at its highest thinking setting, several Argon scores were computed by Google itself, and the other models' numbers come from their makers' own reports. That is the normal way launch tables are built, but it is not a controlled head-to-head. The independent results below are the better guide. 2
What independent testing says
Artificial Analysis runs every model through the same ten evaluations. On its Intelligence Index, Gemini 4 Argon at high reasoning scores 53, level with GPT-6 Astra at maximum effort and one point ahead of GPT-6.1 Sol. It is 23 points above Google's previous non-Flash model, Gemini 3.1 Pro Preview. Claude Opus 5.5 and Claude Sonnet 5.5 still sit higher, at 58 and 56. 3
Cost is where the launch discount does the work. At the introductory price, Artificial Analysis puts Argon at $1.99 per index task, about 60% of GPT-6 Astra's $3.26, but nearly three times GPT-6.1 Sol's $0.72. The saving comes from lower token prices, not efficiency: Argon averaged about 62,000 output tokens per task, more than twice Astra's 27,000. When Google's price doubles after the launch period, the same work would cost about $3.98 by simple doubling, slightly more than Astra. 13
| Model | Input | Cached input | Output | Cost per AA task |
|---|---|---|---|---|
| Gemini 4 Argon (launch) | $2.00 | $0.10 | $10.00 | $1.99 |
| Gemini 4 Argon (standard) | $4.00 | 95% off input | $20.00 | about $3.98 |
| GPT-6 Astra | $10.00 | $1.00 | $50.00 | $3.26 |
| GPT-6.1 Sol | $2.00 | $0.10 | $10.00 | $0.72 |
| Claude Opus 5.5 | $4.00 | $0.20 | $20.00 | $5.98 |
The standout: far fewer made-up answers
The most useful number for a business may be this one. On Artificial Analysis's AA-Omniscience test, which asks hard knowledge questions and penalizes confident wrong answers, Argon's hallucination rate is 15%, far below the other leading models: GPT-6 Astra's is 51% and GPT-6.1 Sol's 54%. In practice, when Argon does not know something, it is much more likely to say so than to guess. 3
There is a trade-off. Argon also answers fewer questions correctly: its accuracy on the same test is 50%, against 63% for GPT-6 Astra. Its overall AA-Omniscience score, which balances the two, lands at 42, level with Astra's 43 and GPT-6.1 Sol's 42. So Argon does not know more; it is better at knowing what it does not know. For legal, finance and client-facing work, where a confident wrong answer is expensive, that is often the better trade. 3

Stronger at getting work done
Agentic work, where a model carries out a multi-step task with tools, has historically been a weaker area for Gemini. Artificial Analysis's results show that changing. Argon ranks first on its version of AutomationBench at 77.5%, ahead of Claude Sonnet 5.5 at 71.3%, and scores 57% on Terminal-Bench 4.0, up from 4% for Gemini 3.1 Pro Preview. 3
Google's own examples point the same way. Inside Google, teams of Argon agents analyzed fleet-wide data to find memory optimizations that freed more than 300 TiB of memory, and Argon agents are migrating large C and C++ codebases to the Rust language, with human review before anything reaches production. Google also raised the output limit to 1 million tokens per response, so the model can work through very long reasoning or generate large documents in one go. 1
Security first, by design
Google trained Argon to be highly capable at cybersecurity defense: finding, validating and patching software vulnerabilities on its own. For trusted defenders and Google's internal teams, it will be released without its cyber guardrails. That is exactly why access is limited. Google describes four safeguard areas it is strengthening first: defending against misuse, resisting prompt injection attacks, monitoring the model's reasoning for signs it is overstepping the user's intent, and hardening the test environments. 1
Prompt injection is the one to understand if you plan to connect AI to email, documents or the web. It is when hidden instructions in content the model reads try to hijack what it does. Google says Argon is its most resilient model yet against these attacks. That lowers the risk; it does not remove the need for approval steps on anything that sends, pays or deletes. 1
What this means for your business
| If you are... | What Argon changes | What to do now |
|---|---|---|
| A law, accounting or finance firm | Leading vendor scores on legal and finance agents, and low hallucination | List two research or review tasks to trial when access opens |
| Automating business workflows | Top AutomationBench results on both Google's and AA's versions | Map one workflow now so it is ready to test |
| Mostly writing code | Strong on DeepSWE; behind on FrontierSWE and Terminal-Bench | Keep your current coding model; compare on your own repo later |
| Cost-sensitive at high volume | Launch price is attractive but doubles later; Sol is cheaper per task | Price on the standard rate, not the discount |
| A security team | Designed for defense; Fairwind access for qualifying organizations | Check whether your organization qualifies for Fairwind |
Two cautions for planning. First, the launch discount is temporary and Google has not said when it ends, so any budget should assume the standard $4 and $20 prices. Second, Argon uses a lot of output tokens per task, which the per-token price hides. Test on your own work and compare total cost per finished task, not the price per million tokens. 13
Comparing it with OpenAI's newest mid-priced model? Our DevDay guide covers GPT-6.1 Sol's price and results.
Monolith's take: Gemini 4 Argon puts Google back among the top AI labs, with leading results on legal, finance and business-workflow tests and a hallucination rate far below its rivals. It is not available to most businesses yet, and its price doubles after launch. Prepare the workflows you would trust it with, and judge it on cost per finished task when access opens.
When can I use Gemini 4 Argon?
How much does Gemini 4 Argon cost?
Is Gemini 4 Argon better than GPT-6 Astra?
Read for this feature. The numbers match the markers in the text.
- Google: Gemini 4 Argon, our next era of frontier intelligenceblog.google
- Google DeepMind: Gemini 4 Argon model evaluation methodologydeepmind.google
- Artificial Analysis: Gemini 4 Argonartificialanalysis.ai
- Google DeepMind: Fairwind Programdeepmind.google
Planning which AI models to trust with client work, and on which tasks? We help teams test and choose.
Want to apply this to a project? Here is the related service.
25 AI instructions to try on everyday work
Prompts are instructions you give an AI tool. This PDF includes 25 examples for inquiries, marketing and admin. Adapt them to your task and check the results.