Google (GOOGL) says Gemini 4 Argon leads most benchmarks. Independent rankings tell a more mixed story

Google, ChatGPT and Claude logos side by side

Key points

  • Google's own table puts Argon on top
  • Independent rankings are more mixed
  • It costs less per task than most top rivals

Alphabet's (GOOGL) Gemini 4 Argon had the top score on 13 of 19 results in the comparison table Google published with the model on Wednesday. Independent evaluations give it a more mixed position, though they use different tests and ranking methods.

On the Artificial Analysis Intelligence Index, Argon scores 53, tied with OpenAI's GPT-6 Astra and Anthropic's Claude Fable 5.1 and behind two other Anthropic models. On Arena's Agent Arena leaderboard, which rates models working as agents, Argon ranked 10th on Friday. On both, Argon costs less per task than GPT-6 Astra and Anthropic's top models, though OpenAI's GPT-6.1 Sol costs less still.

How does Gemini 4 Argon rank on Artificial Analysis?

Artificial Analysis runs every model through the same set of tests and combines them into one index score. The table below shows each model at its highest-scoring setting on the leaderboard as of Friday, along with what Artificial Analysis measured it costs to run one task in the index.

ModelCompanyIntelligence IndexCost per task
Claude Opus 5.5Anthropic58$5.98
Claude Sonnet 5.5Anthropic56$7.67
Claude Fable 5.1Anthropic53$7.63
GPT-6 AstraOpenAI53$3.26
Gemini 4 ArgonGoogle53$1.99
GPT-6.1 SolOpenAI52$0.72

Source: Artificial Analysis, as of Oct. 2, 2026. Scores and costs are live and can change.

Argon's $1.99 cost reflects Google's introductory prices, which are half its standard rates. It also used about 62,000 output tokens per index task, compared with about 27,000 for GPT-6 Astra, R&D World reported. At standard prices, R&D World put Argon's cost at $3.98 per task, more than GPT-6 Astra's $3.26. Google hasn't said when the introductory pricing ends.

How does Argon rank on Arena's Agent Arena?

Arena says Agent Arena ranks models on "how well they orchestrate tools for real-world agentic tasks, based on signals like tool reliability, task completion, and steerability." Argon ranked 10th of 50 models on Friday, based on 10,030 sessions. Arena gave it a possible rank range of fourth to 14th, which reflects how uncertain the ranking still is.

Claude Fable 5.1 ranked first, followed by Claude Opus 5.5, Claude Sonnet 5.5, GPT-6 Astra, and GPT-6.1 Sol. Argon's typical cost per task was $0.67, according to Arena. That's below four of those five models. GPT-6.1 Sol's was $0.56. These costs cover Arena's tasks and shouldn't be compared directly with Artificial Analysis's per-task figures.

Argon did rank first of the 50 on steerability, which Arena describes as "how well the model lands user corrections when they push back."

The ranking has already moved. R&D World reported that Argon was eighth on Oct. 1, based on 3,417 sessions.

Where did Google's own table differ?

Google compiled the comparison table using its own evaluations and some results from external benchmark leaderboards. Argon led on 13 of the 19 results and tied GPT-6 Astra on one. GPT-6 Astra led on three, including the FrontierSWE v2 coding test, and Claude Opus 5.5 led on two.

Google didn't run every competitor score itself. For Terminal-Bench Science, Google tested Argon and took the other models' scores from that benchmark's leaderboard, according to R&D World's reading of Google's methodology.

Bloomberg has reported that Argon does well on benchmarks but less well when Google staff put it to work, and that it struggles with some coding tasks. Google disputed that characterization, according to R&D World.

What does Gemini 4 Argon cost?

Argon's introductory prices are $2 per million input tokens and $10 per million output tokens, and its standard prices are $4 and $20. GPT-6 Astra and Claude Fable 5.1 both list at $10 and $50, according to Arena's leaderboard. Argon is available first to cybersecurity defenders in Google's Fairwind program, and Google hasn't given a date for wider release.

Alphabet shares were up 1.6% at $343.66 as of 12:51 p.m. ET Friday.

Frequently asked questions

How does Gemini 4 Argon rank on Artificial Analysis?

As of October 2, 2026, Gemini 4 Argon scored 53 on the Artificial Analysis Intelligence Index, tied with OpenAI's GPT-6 Astra and Anthropic's Claude Fable 5.1, and behind Anthropic's Claude Opus 5.5 at 58 and Claude Sonnet 5.5 at 56. The leaderboard is live and can change.

Where does Gemini 4 Argon rank on Arena's Agent Arena?

Gemini 4 Argon ranked 10th on Arena's Agent Arena leaderboard on October 2, 2026, based on 10,030 sessions, with a possible rank range of fourth to 14th. Claude Fable 5.1, Claude Opus 5.5, Claude Sonnet 5.5, GPT-6 Astra and GPT-6.1 Sol ranked first through fifth.

Is Gemini 4 Argon cheaper than GPT-6 Astra?

At Google's introductory prices, yes. Artificial Analysis measured Argon at $1.99 per index task, versus $3.26 for GPT-6 Astra. Argon used more output tokens per task, and R&D World put its cost at $3.98 per task at Google's standard prices.

How much does Gemini 4 Argon cost?

Gemini 4 Argon's introductory API prices are $2 per million input tokens and $10 per million output tokens, half its standard rates of $4 and $20. Google has not said when the introductory pricing ends.

Did Google's benchmarks show Gemini 4 Argon in the lead?

In the comparison table Google published on September 30, 2026, Argon had the top score on 13 of 19 results and tied GPT-6 Astra on one. Google compiled the table using its own evaluations and some results from external benchmark leaderboards.

More on GOOGL

Dennis Singleton
Dennis Singleton

Dennis Singleton was born in Australia and later moved to the United States. He has spent years following the markets, but what keeps his attention is how AI is built. He writes about the companies behind the technology, from semiconductor designers and advanced packaging to photonics, memory, networking, and the hardware powering modern AI. His approach starts with filings, earnings, and industry research, then translates the important details into clear, straightforward analysis without unnecessary hype.