Gemini 4 Argon is Google’s new frontier model for extended professional work. Announced on September 30, 2026, it emphasizes coding, enterprise tasks and defensive cybersecurity. The key questions for readers are what the larger output allowance means, what access is available and how to evaluate the announced pricing.
Watch our Gemini 4 Argon overview
The short introduces the announcement. The explanation below separates launch claims from practical evaluation advice. Cover illustration: AI-generated for SoftReviewed.
Gemini 4 Argon at a glance
Three distinctions to understand before planning an integration.
01 · CAPACITY
1M
Output tokens
An output ceiling, not a claim about input context or correctness.
02 · PRICING
$2 in / $10 out
Introductory rates
Per million tokens. Later: $4 input / $20 output. Budget for completed tasks.
03 · AVAILABILITY
Phased access
Check your account
Trusted defenders first. Broader access planned; no general-release date specified.
Access → Representative task → Correctness checks → Actual cost → Adoption decision
Launch facts: Google’s announcement. Workflow advice: SoftReviewed. Recheck availability and billing before adoption.

What is Gemini 4 Argon, in plain English?

Think of Argon as a general-purpose AI model that can support an agent working through a complicated assignment. The model generates responses; an application around it supplies documents, tools, permissions and a way to check the work. It is not a finished accounting service, an automatic deployment system or a guarantee of expert judgment.
The practical distinction is between producing an answer and completing a verified workflow. For example, drafting a migration plan is different from applying that migration safely to a production repository. Argon’s capabilities need to be paired with the right tools and review process. The examples below are proposed evaluation tasks informed by the published results, not demonstrations independently run by SoftReviewed.
When to try Argon—and what to ask it to produce
Repository maintenance
Give it: a small repository, a failing test and the acceptance criteria.
Ask for: a patch, an explanation of the change and test results.
Why evaluate it: the launch chart supports a coding-workflow shortlist, while also showing task-dependent weaknesses.
Pass check: tests pass, unrelated behavior stays intact and a reviewer accepts the diff.
Evidence-backed document work
Give it: a defined collection of reports, filings or policy documents.
Ask for: a structured brief with source references, calculations and unresolved questions.
Why evaluate it: knowledge-work results are strong relative to the displayed competitors.
Pass check: each material claim traces to a source; totals and quotations are checked.
Long, coordinated deliverables
Give it: a specification, shared terminology and a list of required artifacts.
Ask for: a consistent set of modules, documentation or a multi-section report.
Why evaluate it: extended output is relevant when the artifact itself is large.
Pass check: references, interfaces and requirements remain consistent across sections.
Charts and recorded material
Give it: supported visual material and a specific question.
Ask for: extracted observations with the relevant chart region or video time.
Why evaluate it: the supplied chart includes favorable visual-understanding results.
Pass check: inspect axes, units and cited moments; do not accept invented observations.
A specialized use case: authorized security review
For a team with appropriate access and written authorization, a bounded security-review task could involve identifying a suspected flaw, proposing a fix and producing a reproducible validation case. Keep the exercise within approved systems. A model finding is a lead until it is validated; remediation needs regression checks. Argon’s rollout does not make this a generally available consumer workflow.
When not to choose Argon by default
| Situation | Better starting choice | Reason |
|---|---|---|
| Routine sorting, routing or extraction | Rules, a smaller model or a dedicated classifier | Start with the simplest system that meets measured accuracy and latency needs. |
| Immediate production deployment | An accessible, tested model with a fallback | Confirm entitlement and operational limits before depending on a phased rollout. |
| A terminal-heavy coding workflow | Compare Argon with the stronger terminal results in the same evaluation | The supplied chart does not show Argon leading that row. |
| Scientific terminal or post-training work | Task-specific evaluation against the leading candidates | Argon does not lead those displayed rows either. |
| A legal, financial or security decision without review | A qualified reviewer with reproducible evidence | Benchmark performance does not establish professional accountability. |
| A short answer with a strict response-time budget | A model with verified latency on your workload | Large output capacity is not a speed guarantee; Argon speed data was unavailable on the checked model page. |
Potentially weaker fit does not mean incapable. These are selection recommendations based on the published comparisons and workflow requirements. Test the actual assignment before deciding. For simple tasks, avoid paying for complexity you do not need; for difficult tasks, measure correctness and total completion cost rather than response length.
A simple decision path
Can your account access it? If no, use an available alternative.
Does the assignment need extended reasoning or a substantial artifact? If no, test a simpler route first.
Can you define an objective success check? If no, add a reviewer and clarify the requirements.
Does Argon pass representative tasks at an acceptable cost? If yes, introduce it gradually with a fallback.
What does one million output tokens mean?
Google specifies a maximum output allowance of one million tokens, compared with the previous 64,000. This figure should not be presented as an input-context specification.
For a developer, output capacity and input capacity answer different questions. Input capacity concerns the material a system can receive; output capacity concerns what it can produce. Neither number alone demonstrates that a generated migration, report or application is correct.
Our practical recommendation is to evaluate completion rather than length. Give a model a representative task with clear acceptance criteria, inspect the artifact, and measure how much human correction it needs. A long response that fails tests is less useful than a shorter response that satisfies the requirements.
Pricing: introductory rates versus later rates
| Announced rate per million tokens | Input | Output |
|---|---|---|
| Introductory | $2 | $10 |
| After the introductory period | $4 | $20 |
Google also announces a 95% cached-input discount. These are announced launch terms, not confirmation that every reader can buy access today.
Budget by completed task
Before choosing an API for production, record the total input, output, retries and tool activity required for a successful task. Check the actual billing documentation when access opens: caching eligibility, reasoning-token treatment and other charges can matter. Do not budget by multiplying a headline rate without confirming how the service meters the workload.
Who can access Gemini 4 Argon?
Initial access is limited to trusted cyber defenders through Fairwind. Google plans broader availability, starting with paid API customers and Google AI Ultra subscribers; the announcement does not specify a general-release date.
For readers planning an integration, treat an announcement and a deployable endpoint as separate milestones. Before changing production routing, confirm your account can access the exact model, review its limits and test a fallback. A subscription label alone should not be treated as evidence that the model is already enabled.
How should developers evaluate the engineering claims?
Google reports a 2.7× improvement for a Rust video-decoder port while preserving output. This is a particular engineering result, not a universal acceleration guarantee.
A useful coding evaluation needs more than an impressive demo. Select a real maintenance task, hold the repository and requirements constant, and compare the resulting code against your existing workflow. Run functional tests, inspect regressions, and record reviewer time. For performance-sensitive code, benchmark on the same hardware with the same inputs.
Keep approval boundaries clear for an agent that can edit files or call tools. Start with an isolated branch or disposable environment, require review before deployment, and preserve logs of the actions taken. These are our workflow recommendations, not claims that we independently tested Argon.
What should you check before adopting it?
- Availability: confirm the model and region are actually enabled for your account.
- Quality: define correctness checks before comparing outputs.
- Cost: compare successful-task cost, including failed attempts.
- Reliability: inspect error handling, tool permissions and recovery behavior.
- Privacy: review the applicable service terms before sending company documents.
For an example of evaluating a different kind of AI workload, see our guide to decision engines and traditional language models. The useful comparison is whether a tool fits your task, rather than whether it wins every category.
Benchmark comparison: where Argon leads—and where it does not
The launch comparison supplied with Google’s announcement contains four models and 19 scored rows. The table below transcribes that chart, including its subset labels. These are Google-reported results, not SoftReviewed tests. A dash means no result is shown; it is not zero. Blue identifies Argon, while bold identifies the highest reported score in each row, including ties.
| Area | Benchmark | Gemini 4 Argon | GPT-6 Astra | Claude Fable 5.1 | Claude Opus 5.5 |
|---|---|---|---|---|---|
| Knowledge work | Vals Index | 68.9% | 63.1% | 65.8% | 67.0% |
| Knowledge work | AutomationBench | 51.3% | 41.4% | 31.4% | 42.5% |
| Knowledge work | Vals Finance Agent v2 | 65.4% | 53.5% | 58.9% | 58.6% |
| Knowledge work | Harvey’s Legal Agent Benchmark | 19.6% | 5.4% | 6.7% | 3.8% |
| Agentic coding | DeepSWE v1.1 | 77.9% | 74.1% | 67.4% | 74.2% |
| Agentic coding | FrontierSWE v2 | 55.0% | 65.5% | 56.3% | 62.3% |
| Agentic coding | Vibe Code Bench | 91.9% | 89.6% | 90.3% | 90.3% |
| Agentic coding | Terminal-bench 4.0 | 57.4% | 58.2% | 57.9% | 66.4% |
| ML engineering | PostTrainBench | 45.3% | 44.3% | 40.2% | 49.3% |
| Science and math | Terminal-Bench Science 0.1 | 57.6% | 68.1% | 52.6% | 63.3% |
| Science and math | LABBench 2 | 88.8% | 85.4% | 68.6% | 73.1% |
| Science and math | RiemannBench | 76.0% | 72.0% | 65.6% | 69.6% |
| Long context | GraphWalks: up to 128k, BFS (F1) | 99.7% | 98.7% | 91.4% | 90.6% |
| Long context | GraphWalks: 256k–1M, BFS (F1) | 84.2% | 71.8% | 65.0% | 66.8% |
| Computer use | Agent’s Last Exam: pass rate | 39.5% | 34.2% | — | 38.2% |
| Computer use | OSWorld-2.0: offline subset, partial score | 69.2% | 72.6% | — | — |
| Multimodal understanding | Chartography | 71.6% | 71.0% | 46.2% | 66.3% |
| Multimodal understanding | LVBench | 91.7% | 87.5% | 79.7% | 83.7% |
| Cybersecurity | CWE-bench v1 | 68.0% | 68.0% | 58.0% | 67.0% |
Methodology reference: Google’s Argon evaluation methodology. This methodology page could not be retrieved during our verification; reasoning budgets, harnesses and repeated-run details remain unverified here. Do not assume these scores use the same setup as another evaluator.
Read the results by workload
Knowledge work
Argon leads all four listed knowledge-work rows. That makes document-heavy work a sensible evaluation target. It does not remove the need to check citations, calculations and whether the answer actually follows the brief.
Coding is not one skill
Argon leads DeepSWE and Vibe Code Bench; Astra leads FrontierSWE and Opus leads the terminal row. A claim that one model is simply “best at coding” would hide these differences. Choose a test resembling your repository and tooling.
Long context and visual work
The graph and multimodal rows favor Argon in this chart. Treat graph traversal, chart interpretation and video understanding as separate capabilities. Passing one does not establish reliable performance on every document or diagram.
Important limits
Argon does not lead the displayed post-training or scientific terminal tasks. Cybersecurity is a tie with Astra. The computer-use subset is also mixed. These are reasons to compare by task, rather than declare a universal winner.
Why a strong relative score can still need human review
First place in a row is a relative result among the models shown. It is not the same as a high absolute success rate. A model can beat its peers while still failing many examples. Percentages from different rows also cannot be averaged meaningfully without accounting for task definitions, sampling and the evaluator’s weighting.
For document work, keep the original material available to a reviewer. For code, require tests and inspection of the patch. For desktop automation, confirm the final application state. For security work, require authorized scope and a reproducible validation process. The point of the benchmark is to help select a candidate for evaluation, rather than to replace those controls.
Argon versus 11 other models: independent index and cost
Snapshot: October 1, 2026. One displayed configuration per model; reasoning settings differ. Index points are not percentages. Task cost is an evaluation-specific measurement, not API token pricing or a quote for your workload. Argon uses introductory pricing.
| Model/configuration | Intelligence Index | Cost per index task |
|---|---|---|
| Gemini 4 Argon (high) | 53 | $1.99 |
| Claude Opus 5.5 (max with fallback) | 58 | $5.98 |
| Claude Sonnet 5.5 (max with fallback) | 56 | $7.62 |
| Claude Fable 5.1 (max with fallback) | 53 | $7.63 |
| GPT-6 Astra (max) | 53 | $3.26 |
| GPT-6.1 Sol (max) | 52 | $0.72 |
| Muse Spark 1.3 (max) | 48 | $1.60 |
| Grok 4.7 (xhigh) | 46 | $3.74 |
| MiMo-V2.6-Pro | 46 | $0.13 |
| Qwen3.8 Max (0902) | 45 | $5.41 |
| GLM-5.3 (max) | 45 | $2.01 |
| Kimi K3 (max) | 44 | $2.00 |
Source: Artificial Analysis model leaderboard. This is a selected comparison set, not a claim that these are the top 12 entries or that tied scores mean identical capabilities. Check the live source before procurement.
Keep independent and vendor evaluations separate
Artificial Analysis reports about 78% for Argon on AutomationBench-AA and about 57% on Terminal-Bench 4. Its launch analysis also reports a 15% hallucination rate and 50% accuracy on AA-Omniscience. These are distinct measurements; the low hallucination rate does not imply 85% general accuracy. Source: Artificial Analysis’s Argon launch analysis.
In particular, AutomationBench-AA is not the same reported evaluation as the AutomationBench row in Google’s chart. Do not merge their scores into one ranking. Similarly, terminal scores can differ with the evaluator’s model configuration and harness. Keep the result attached to its source instead of choosing whichever number makes a model look strongest.
A practical evaluation plan for your own work
1. Define success before prompting
Pick a small collection of representative tasks, not only a demonstration designed to look impressive. Include an ordinary task, a difficult task and a case where the correct behavior is to ask for clarification or acknowledge missing information. Write down what counts as success before seeing any response.
2. Compare the same deliverables
Provide equivalent source material and required outputs. Record model version, reasoning configuration, tools and time budget. Allow necessary platform differences, but document them. A comparison with unrestricted tools on one side and no tools on the other answers a different question from a controlled model comparison.
3. Count correction and recovery costs
Track failed runs, reviewer minutes, fixes and whether a task can resume after interruption. A polished first answer can hide expensive cleanup. For an agent, record whether it took the right action in the destination system, not merely whether it described that action convincingly.
4. Adopt gradually
Start with a reversible workflow and keep your existing route available. Expand only after the task-specific checks pass consistently. This provides a more useful adoption decision than choosing a model solely by a headline score.
Frequently asked questions
Does a higher index score mean a model is better for every job?
No. A composite index helps summarize results across its evaluation mix. Your task may emphasize a different capability, and nearby scores should not be treated as certainty about an individual outcome.
Should I select the cheapest model in the comparison?
Use the table to shortlist candidates. Then measure whether the cheaper route finishes your actual tasks with acceptable quality and review effort. A lower evaluation cost is useful evidence, but it does not guarantee the cheapest complete workflow.
What does a missing benchmark result mean?
It means the supplied chart does not show a result for that model on that row. It does not establish that the model cannot perform the task, and it should never be converted to zero.
Why are there different terminal and automation scores?
Evaluators can use different task sets, configurations and execution harnesses. Keep version and source labels visible and compare results within the same evaluation setup.
Is SoftReviewed claiming hands-on testing?
No. This article combines attributed published results with our interpretation and evaluation recommendations. We have not run an independent Argon test suite. The local helper’s synthesized benchmark fields were excluded from the comparison.
Sources and reporting limits
Launch facts are attributed to Google’s official Gemini 4 Argon announcement. SoftReviewed reviewed early commentary to identify reader questions, but has not independently tested the model. Benchmark tables above cite the supplied launch chart and independently retrieved Artificial Analysis data; speculative release dates from commentary are excluded. Access and billing details should be rechecked when the rollout changes.







