MATINTELLECT

👀 After months of delays Google finally rolled out the flagship of the Gemini 4 generation - the Argon model. By the company's own table it beats Opus 5.5, GPT-6 Astra and Fable 5.1. There's just one catch: almost nobody can get their hands on it yet

What the Argon announcement showed:
The September 30 post was written by Koray Kavukcuoglu, SVP at Google DeepMind and the company's chief AI architect. Argon is built for long, complex tasks: real-world development, finance, legal work and cybersecurity. Out of 19 benchmarks in Google's table the model comes first in 13. Sounds like a rout, but I dug deeper into the numbers

Where Argon is genuinely ahead:
🟢 DeepSWE v1.1 - 77.9% vs 74.2% for Opus 5.5 and 74.1% for Astra. Long multi-step development tasks, the loudest result of the release
🟢 Vals Index - 68.9% vs 67.0% for Opus 5.5. Finance, law, taxes and code in one score
🟢 GraphWalks at 1M tokens - 84.2% vs 71.8% for Astra. Working with huge context
🟢 AutomationBench - 51.3%, first place

And where it loses - and that's more interesting:
🟠 Terminal-Bench 4.0 - 57.4% vs 66.4% for Opus 5.5. Almost a 9-point gap, and that's exactly the test about agentic work in the terminal
🟠 FrontierSWE v2 - 55.0% vs 65.5% for Astra and 62.3% for Opus 5.5

The picture is a weird one: in one coding benchmark Argon leads, in the other two it lags noticeably. All the numbers come from Google itself, there's no independent verification yet. How much of these percentages will carry over into real work only becomes clear once the model is in developers' hands

The feature nobody's talking about much: Argon outputs up to 1M tokens in a single response. The ceiling used to be 64K. This isn't about "writing a long essay", it's about being able to generate an entire module or migration in one go, without chopping it into pieces

Right now only vetted cybersecurity specialists through the Fairwind program and Google's internal teams have access. For them the model runs without cyber restrictions: it finds, verifies and patches vulnerabilities on its own. Google says outright that it held back the release so Argon wouldn't fall into hackers' hands. There's no public ID, and it's not in Vertex AI, Gemini CLI, Cursor or Copilot yet either

💎 Price: $2/$10 per 1M tokens at launch, then $4/$20

The launch price is half that of Opus 5.5, plus a 95% discount on cached input. Google AI Ultra subscribers and the paid API get the model first. No dates, and how long the promo price will last isn't said either

⭐️ Koray Kavukcuoglu, SVP Google DeepMind:

"Argon fundamentally changes the way we work and build products at Google"

🤷 Funny math: after the promo Argon will cost exactly the same as Opus 5.5. So Google is dumping prices for exactly the stretch when everyone's comparing benchmarks. For now I'm staying on Opus and I see no point in switching in the terminal - Argon loses there by 9 points. But if DeepSWE isn't lying about long tasks, the very first month of open API will settle it

Instagram | YouTube | Threads

Share:

No comments yet

Leave a Comment

Fields marked with an asterisk (*) are required