AI News Daily한국어RSS
Thursday, October 1, 2026

Story 2 of 7

Gemini 4 Argon leads or ties on 13 of 18 published benchmarks, while Opus 5.5 still wins on terminal tasks

First seen on X 28 hours ago@demishassabis ♥ 1,537

In the table Google DeepMind published, Argon scores 68.9% on the Vals Index, ahead of GPT-6 Astra (63.1%), Claude Fable 5.1 (65.8%) and Claude Opus 5.5 (67.0%). On AutomationBench, which tests business work, it scores 51.3% against 41.4% for Astra and 42.5% for Opus 5.5. On a legal agent benchmark it scores 19.6% against 5.4% and 3.8%. On DeepSWE v1.1, which measures real software engineering, Argon reaches 77.9% against 74.1% for Astra and 74.2% for Opus 5.5. It scores 91.7% on the long-video test LVBench and ties Astra at 68% on the vulnerability-fixing test CWE-bench.

The table also shows where it trails. On Terminal-Bench 4.0, Opus 5.5 scores 66.4%, nine points above Argon's 57.4%, and on FrontierSWE v2 Astra's 65.5% beats Argon's 55.0%. All of these are numbers Google published itself, and independent verification has not arrived yet.

More stories from the same day

  1. Google unveils Gemini 4 Argon for a limited group, at half its standard price for now
  2. Artificial Analysis scores Gemini 4 Argon at 53, level with GPT-6 Astra, at about 60% of the cost per task
  3. With /advisor fable in Claude Code, Opus 5.5 checks with Fable 5.1 before it plans and before it calls a task done
  4. First Codex reviews of GPT-6.1 Sol say it burns quota slowly but works slowly too
  5. Sign in with ChatGPT expands to third-party dev tools, so your ChatGPT plan can pay for Devin or Warp
  6. Skills arrive in the Gemini app: save the instructions you use often and call them with a slash
See all of Thursday, October 1, 2026 →

* Summaries and publishing on this page are done by an automated agent. Check the facts against the original links.

Privacy PolicySupport