NewsStocksxAI Launches Grok 4.7 After Repeated Delays, Trails Frontier Rivals on Benchmarks

xAI Launches Grok 4.7 After Repeated Delays, Trails Frontier Rivals on Benchmarks

Author: Decrypt·

Key Takeaways

  • Grok 4.7 uses 2.1 trillion parameters, a 40% increase over the 1.5 trillion in Grok 4.6, and is priced at $2 per million input tokens and $6 per million output tokens.
  • xAI supplemented the model's training with SpaceX data, includinglink satellite telemetry, manufacturing records, and engineering failure logs, aiming to improve reasoning about hardware and physical systems.
  • Early benchmark results rank Grok 4.7 second to Claude Fable 5.1 on GDPval (1695 vs. 1735) and AA-Briefcase (1657 vs. 1678), and second to GPT-6 Astra on the EEBench electrical-engineering test.
  • The launch followed at least five revised release timelines since late July, and the model went live immediately without a waitlist in the Grok app, Cursor, Grok Build, and the xAI API.
  • Musk said Grok 4.7 should land roughly on par with Anthropic's Claude Opus 5.0, and sketched upcoming models—Grok 4.8, Grok 4.9, and a possible frontier-leading Grok 5—without assigning release dates.
xAI Launches Grok 4.7 After Repeated Delays, Trails Frontier Rivals on Benchmarks

Elon Musk's xAI released Grok 4.7 on Monday afternoon, its most capable model to date, which the company called "a notable improvement over Grok 4.6 at the same price and speed."

Early benchmark results placed the model second to Claude Fable 5.1 on GDPval and AA-Briefcase, and second to GPT-6 Astra on EEBench, an electrical-engineering benchmark.

The launch arrived after a string of slipped deadlines. Musk walked back the release timeline at least five times since late July: first "four weeks out," then "a few weeks," then "3 to 4 weeks," then "10 days" on September 1, and finally "needs a few more days to cook" on September 11.

There is no waitlist this time. Grok 4.7 is live immediately in the Grok app, Cursor, Grok Build, and the xAI API.

Musk followed up on X, calling Grok 4.7 "a strong combination of intelligence, speed & low cost."

Grok 4.7 is here. It's a notable improvement over Grok 4.6 at the same price and speed. pic.twitter.com/H3OTBbXyvO

SpaceXAI (@SpaceXAI) September 21, 2026

According to xAI's official announcement, the model spends longer working through hard problems and double-checks its own answers more often than Grok 4.6 did, alongside what the company calls its strongest safety guardrails yet.

Scale, Pricing, and SpaceX Data

Grok 4.7 packs 2.1 trillion parameters, up 40% from the 1.5 trillion in Grok 4.6, which was itself a refinement of Grok 4.5. Costs are set at $2 per million input tokens and6 per million output tokens. Parameters are the internal knobs a model tunes during training, and more of them generally means more capacity to learn patterns, while tokens are the basic units of information an AI model can either register or generate.

xAI also folded in supplemental training data pulled from SpaceX, Musk's rocket company: Starlink satellite telemetry, manufacturing records, and engineering failure logs. The pitch is a model that reasons better about hardware and physical systems than anything trained purely on internet text.

Benchmarks Tell a Familiar Story

GDPval measures how a model performs on real, economically valuable knowledge work—legal memos, spreadsheets, slide decks—using tasks vetted by working professionals in each field, and scores it as an Elo rating, the same head-to-head ranking system chess uses. Grok 4.7 hit 1695 on GDPval, while Claude Fable 5.1 topped the chart at 1735—a 40-point gap on the Elo scale.

AA-Briefcase, built by Artificial Analysis, tests multi-hour office work that strings research, analysis, and document production into one long task, also scored on the Elo scale. Grok 4.7 posted 1657 against Fable 5.1's 1678—a narrower 21-point gap, and the same result on a different test.

CursorBench 4.0, Cursor's benchmark for real coding tasks inside its editor, plots accuracy against the cost and token count each task burns through. Grok 4.7 lands in the middle: pricier per task than GPT-6 Astra, the successor to GPT-5.6 Sol, and Claude Sonnet 5, but still short of Fable 5.1, which wins at every price point on the chart.

A Recurring Pattern for xAI

The company has repeatedly launched flagship models that arrived behind rivals on independent evaluations. Grok 4.5 debuted in July with the biggest training cluster in the industry and third-place scores behind Claude and OpenAI's models. Before it, Grok 4.20 traded reliability for speed and personality, and Grok 4.6 trailed the frontier pack on coding autonomy.

None of that makes Grok 4.7 a bad product for the millions of people who talk to it through X, the standalone app, or their Tesla's dashboard. It does mean the model most likely to answer users' questions—or power their car's voice assistant—runs on a system that its own maker's benchmarks place a rung below the top of the ladder.

That gap is why the price tag matters more than the leaderboard position for most people. xAI has consistently undercut Anthropic and OpenAI on cost per token even as it trails them on raw capability, betting that "good enough, cheap, and everywhere" beats "best, but pricier" for the bulk of everyday use.

Musk's Expectations and the Road Ahead

Musk had already set expectations lower days before launch, writing on X that Grok 4.7 should land "roughly on par with" Anthropic's Claude Opus 5.0—not the newer Opus 5.1—with multimodal performance still needing work.

In that same post, he sketched the next three models: Grok 4.8 as a meaningful step up, Grok 4.9 in "Astra/Fable class," and Grok 5 as a possible frontier leader. Those upgrades do not yet have release dates—a detail that carries extra weight given Grok 4.7's own timeline was walked back at least five times before launch. "We shall see," he wrote.