Gemini 4 Argon Is Here (and Nobody Can Use It) — A Deep Dive Into Google’s New Frontier Model

Gemini 4 Argon Is Here (and Nobody Can Use It) — A Deep Dive Into Google’s New Frontier Model

Published October 1, 2026

Google just unveiled Gemini 4 Argon, its most powerful model to date. The announcement came quietly through a blog post by Koray Kavukcuoglu, who leads Google DeepMind — but the model itself has been hiding in plain sight for weeks under fake names on the LMSYS Arena, where blind evaluators kept spotting it and rating it among the best.

The price is $2 per million input tokens and $10 per million output tokens. It can output up to 1 million tokens in a single response. It has the lowest hallucination rate Artificial Analysis has ever measured on a model scoring above 40 intelligence points: just 15%. And Google is already using it internally to rewrite its own infrastructure, with one notable result — a Rust-based video decoder that runs 2.7 times faster than the previous version.

Almost nobody outside Google can use it yet. The model is currently restricted to a program called Fairwind, reserved for trusted cyber defenders. But I got my hands on a game that an early version of Gemini 4 built — and I played it. The results, compared to every other frontier model, tell the real story.


1. The Announcement: What Google Actually Said

On September 30, 2026, Koray Kavukcuoglu published the official announcement. Gemini 4 Argon is Google DeepMind’s new frontier model. For now, access is limited to vetted cybersecurity professionals through a program called Fairwind. The reasoning is cybersecurity — Argon can independently find, validate, and patch software vulnerabilities. Google wants to monitor the model’s reasoning for misalignment before opening it broadly.

The pricing structure is aggressive:

  • $2 per million input tokens (introductory — doubles to $4.20 after the promo period)
  • $10 per million output tokens
  • Cached input is 9× cheaper, effectively a 95% discount
  • 1 million token output limit — a massive jump from the previous 64K ceiling on Gemini models

The model can produce hundreds of thousands of reasoning tokens in a single run. This alone changes the game for long-form analysis, code generation, and agentic workflows.


2. Argon Was Already Winning on the Arena

Before the official announcement, Gemini 4 was competing on the LMSYS Arena under anonymous checkpoints. Blind evaluators kept detecting it — and the scores were remarkable.

LuminaBench (checkpoint “Colosseum”) noted that the style had improved dramatically since earlier checkpoints. AI Screening gave Gemini 4 and Claude Opus 5.5 the same task: build a 3D hovering airplane with realistic physics. The verdict was clear: “Gemini completely surpasses Opus here.” Gemini finished the task in 21 minutes — significantly faster than any other model tested.

Another evaluator asked for a mechanical bee. The result? Intricate wings and gears, fully interactive in the browser. A few days earlier, the same evaluator had obtained a mechanical butterfly of similar quality.

Abhinav obtained a Halloween diorama from the September 28 checkpoint — a ghost with a lantern, pumpkins, candles — all running in the browser. But Gemini 4 doesn’t win at everything. When Abhinav gave the exact same instruction to four models — Sonnet 5.5, Gemini 4, GPT-6.1 Soul, and Opus 5.5 — the evaluators chose Sonnet 5.5 for best details and textures. Gemini only won in the moonlight category.

SPAC89 noticed something different about Gemini 4: it teaches. Ask it how a factory works, and it creates an interactive lesson for you. No other model does it this way, and it’s extremely fast.


3. Internal Use: Argon Is Already Rewriting Google

This is the part most coverage omits: Google is already using Argon internally, and the results are staggering.

  • Quantum computing: Argon surpassed a published quantum computing benchmark by 40% in minutes.
  • Data center optimization: A team of Argon agents reviewed Google’s data centers and freed up more than 300 terabytes of memory.
  • C++ to Rust migration: Google is migrating all C and C++ code to Rust across the company. One example: up to 800,000 lines for a single kernel. Argon is doing the bulk of the translation.
  • Video decoder: In Google’s video decoder, Argon replaced 32,000 lines of hand-tuned code. The new Rust version is 2.7 times faster than the previous port.

Is this AI improving itself? There’s no evidence of that directly — but an AI that accelerates the engineers who build the next AI is exactly how that flywheel starts.


4. Benchmark Dominance — and Weaknesses

Where Argon Leads

  • DeepSWE: 77.9% — a new record for real-world, long-duration software tasks
  • VALS Index: Number one across finance, legal, tax, and coding work
  • AutomationBench (Banking): 51.3% — top ranking
  • Long Video Understanding (LV Bench): 91.7% — better than anything else
  • Legal Reasoning: 19.6% accuracy — a massive gap over competitors, who all score below 7%
  • Long Context (256K–1M tokens): Argon scores 84, GPT-6 Astra scores 72

Where Argon Loses

  • Web of Frontier: Astra scores 65.5, Argon scores 50
  • Terminal Bench: Opus 5.5 is clearly ahead
  • OS World (Computer Use): Astra wins again

Google’s own scorecard is honest about these gaps. Argon is not a universal winner — it excels in specific domains while lagging in others.


5. Independent Verification: Artificial Analysis

Artificial Analysis gives Argon 53 points on its intelligence index — exactly the same as GPT-6 Astra, one point above GPT-6.1 Soul, and a massive 23 points above Google’s latest Pro model. The cost per task at the introductory price is $1.99 (Astra costs $3.26).

But there’s a catch: Argon writes a lot. It produces 62,000 tokens per task on average, compared to Astra’s 27,000. At full price, that would make each task $3.98 — more expensive than Astra.

The hallucination rate is real: just 15%, the lowest ever measured for any model above 40 intelligence points. Astra hallucinates 51% of the time. Soul hallucinates 54%. But fewer hallucinations don’t mean Argon knows more. Its accuracy on the same test is 50%, while Astra’s is 63%. Argon simply says “I don’t know” more often instead of fabricating an answer.

According to Reuters, Argon is physically larger than Google’s previous Pro models. Google didn’t just train smarter — it trained on a larger scale.


6. Cybersecurity: Why Fairwind Exists

Argon can independently find, validate, and patch software vulnerabilities. With Wiz, it found a critical flaw that exposed personal data in hospital software used worldwide. Previous frontier models had completely overlooked it.

This capability is the reason for the slow, restricted deployment. Trusted Defenders get immediate access through Fairwind. Everyone else waits while Google monitors the model’s chain-of-thought reasoning for misalignments.


7. The Real Test: Same Prompt, Same Task

I ran the definitive comparison. YouWare gave Gemini 4 and GPT-6 Astra the same long prompt: “A cozy tram-riding game between floating islands with stations, passengers, a comfort meter, and a workshop to upgrade the tram.” Both models published playable games, and I played them both.

GPT-6 Astra’s Version

A clean interface with a small route map in the corner. Hold W to move the tram across the bridge. 12 passengers on board. Open the doors at the station, then enter the workshop to add upgrades to the tram. Functional, organized, works well.

Gemini 4’s Version

The light is different — a golden hour glow, the sea under the clouds, a lighthouse in the distance. There’s a comfort meter: take a curve too fast and passengers will notice. The tram has autopilot that slows down and stops at stations on its own. There’s even a cabin camera for a first-person view. Oliver’s Cloudworks is a small island workshop where you install upgrades.

The visual and atmospheric difference is dramatic.

What You Can Use Today

I gave the exact same prompt to Gemini 3.8 Flash — the best Gemini you can actually use via the API. One attempt, 3 minutes, $0.11. It works. There’s a station, doors open, passengers board, a comfort meter, even the workshop. But look at the camera: you go straight through the trees. The islands are flat green plates. The comfort meter drops to 2% on the first curve. Same company, same prompt. That’s how big the jump to Gemini 4 is.

Then I gave the same prompt to the models you can afford today:

  • Claude Sonnet 5.5: 12 minutes, $0.92. A warm sunset, the tram pulls away like a real one. 14 passengers. Best longevity.
  • GPT-6.1 Sol (GPT-4o successor): Less than 6 minutes, $0.19. A starry night, a long bridge over the clouds, crosswind affecting the comfort meter. Cheapest option.

My verdict: Gemini 4 Argon built the most beautiful and immersive world. GPT-6 Astra built the most organized and functional game. Of the models you can use right now, Sonnet 5.5 has the best gaming longevity and GPT-6.1 Sol is the cheapest by far. But if Argon really does come out at $2, it’s the best deal in AI.


8. Play All Five Games Yourself

All five games — from Gemini 4 Argon, GPT-6 Astra, Gemini 3.8 Flash, Claude Sonnet 5.5, and GPT-6.1 Sol — are available to play right now at the link below. Each was generated from the exact same prompt with a single instruction, no human intervention.

👉 Play all five games at ruocco.it/ai-demos/gemini-4-cloudline/


9. Watch the Full Breakdown

Watch the video where I test all five models live, play each game, and break down every benchmark:


Final Verdict: Has Google Returned to the Top?

Gemini 4 Argon is a genuine leap forward. Its hallucination rate is the lowest we’ve ever seen at this level. Its internal use at Google is producing concrete, measurable improvements. Its benchmark performance in coding, long context, legal reasoning, and video understanding is genuinely class-leading.

But it’s not a universal winner. It loses in computer use and agentic browsing. Its accuracy is lower than Astra’s — it just says “I don’t know” more often. And the restricted deployment through Fairwind means we can’t fully assess it yet.

What’s undeniable is the gap between Argon and the Gemini models we can use today. The same company, the same prompt — and the difference is shocking. If this is what Google can do when it trains at scale without compromise, the next year of AI is going to be very interesting.

Has Google returned to the top? Tell me in the comments. And if this breakdown helped you, like and subscribe — it really helps the channel.

Sources: blog.google (Gemini 4 Argon) · deepmind.google/models/gemini · artificialanalysis.ai · Reuters (via @AiBattle_) · @LuminaBench · @AI_Screening · @thtbee_ · @abhinavflac · @SPAC89 · @YouWareAI · @AiBattle_

Leave a Comment