Gemini 4 Argon leads or ties on 13 of 18 published benchmarks, while Opus 5.5 still wins on terminal tasks
First seen on X 28 hours ago@demishassabis ♥ 1,537
In the table Google DeepMind published, Argon scores 68.9% on the Vals Index, ahead of GPT-6 Astra (63.1%), Claude Fable 5.1 (65.8%) and Claude Opus 5.5 (67.0%). On AutomationBench, which tests business work, it scores 51.3% against 41.4% for Astra and 42.5% for Opus 5.5. On a legal agent benchmark it scores 19.6% against 5.4% and 3.8%. On DeepSWE v1.1, which measures real software engineering, Argon reaches 77.9% against 74.1% for Astra and 74.2% for Opus 5.5. It scores 91.7% on the long-video test LVBench and ties Astra at 68% on the vulnerability-fixing test CWE-bench.
The table also shows where it trails. On Terminal-Bench 4.0, Opus 5.5 scores 66.4%, nine points above Argon's 57.4%, and on FrontierSWE v2 Astra's 65.5% beats Argon's 55.0%. All of these are numbers Google published itself, and independent verification has not arrived yet.