Grok 4.6 Challenges OpenAI and Anthropic’s Top AI Models

grok

SpaceXAI’s Grok 4.6 challenges OpenAI and Anthropic’s top AI models, matching GPT-5.6 Sol on key benchmarks while undercutting them on price.

Grok 4.6 challenges OpenAI and Anthropic’s top AI models

Grok 4.6 challenges OpenAI, Anthropic’s top AI models as SpaceXAI pushes its latest flagship into the same performance tier as GPT-5.6 Sol and Claude Opus 5. The company says the new model delivers frontier-level intelligence across coding, agentic tasks, and advanced knowledge work, while costing significantly less per token than its rivals.

Elon Musk described Grok 4.6 as “objectively #1 when considering intelligence, speed & cost” in a post on X, framing the release as a direct challenge to the two labs where Microsoft holds major AI stakes. Independent benchmarking from Artificial Analysis supports the claim that Grok 4.6 has closed much of the gap with the current leaders.

Benchmark performance

On the Artificial Analysis Intelligence Index, a composite score built from nine separate benchmarks, Grok 4.6 scores 61, tying OpenAI’s GPT-5.6 Sol (max) and trailing only Anthropic’s Claude Opus 5 (63) and Claude Fable 5 (62). That represents a five-point jump over Grok 4.5 and a 23-point improvement over Grok 4.3.

For agentic and knowledge-work tasks, Grok 4.6 achieves a GDPval-AA v2 Elo score of 1,753, placing it second behind Claude Opus 5 and ahead of most other frontier models. On complex workflows, it reportedly completes tasks in around 53 steps, compared to 103 steps for Claude Opus 5, suggesting greater efficiency in multi-step reasoning.

Pricing and cost per task

SpaceXAI has positioned Grok 4.6 as a lower-cost alternative to OpenAI and Anthropic. Pricing is listed at $2 per million input tokens and $6 per million output tokens, compared to:

  • Claude Opus 5: $5 input, $25 output.
  • GPT-5.6 Sol: $5 input, $30 output.

Artificial Analysis estimates the average cost per task at around $0.84 for Grok 4.6, more than 60% below Claude Opus 5 and GPT-5.6 Sol. That combination of frontier performance and lower inference cost could be particularly attractive to developers, startups, and enterprises running large-scale AI workloads.

Coding and agentic capabilities

Grok 4.6 is designed to compete directly on coding and agent-driven workflows. SpaceXAI says the model achieves frontier intelligence across several agentic coding and knowledge-work benchmarks, including:

  • Complex code generation and debugging.
  • Multi-step task automation.
  • Real-world knowledge work on a computer.

Earlier Grok 4 variants already led on SWE-bench Verified, a coding benchmark that matters heavily to developers. Grok 4.6 extends that strength into broader agentic tasks, where the model must plan, execute, and verify multi-step workflows rather than just generate code snippets.

Why this matters

The release arrives at a critical moment for OpenAI and Anthropic, both of which are preparing for potential IPOs and need to demonstrate strong margins and sustained leadership. Grok 4.6’s ability to match or approach their best models at a much lower price puts pressure on their pricing power and long-term positioning.

For users, the competition is beneficial. More models at the frontier means better performance, faster iteration, and more options for balancing cost, speed, and capability. If Grok 4.6’s real-world performance holds up across different workloads, it could become a serious alternative to GPT-5.6 Sol and Claude Opus 5 for coding, automation, and knowledge work.

Summary: SpaceXAI’s Grok 4.6 matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index and trails only Anthropic’s top models, while costing more than 60% less per token. With strong agentic and coding performance, Grok 4.6 challenges OpenAI and Anthropic’s leadership at the frontier.

Read Previous

Google’s Gemini App Hits 1 Billion Monthly Active Users

Read Next

Facebook Launches Standalone Creator Studio App With AI Tools