NewsStocksElon Musk's SpaceXAI Launches Grok 4.7 With 71% DeepSWE Benchmark Score

Elon Musk's SpaceXAI Launches Grok 4.7 With 71% DeepSWE Benchmark Score

Author: Coinotag·

Key Takeaways

  • SpaceXAI released Grok 4.7 on September 21, describing it as its most capable model yet for coding and knowledge work, with training focused on tasks that take several hours to complete.
  • The model launches at $2 per million input tokens and $6 per million output tokens, identical to Grok 4.6 and far below GPT-5.6 Sol at $4/$20 and Fable 5.1 at $10/$50.
  • On SpaceXAI's published benchmarks, Grok 4.7 scored 71.0% on DeepSWE v1.1, raised Terminal-Bench 4.0 to 38.0% from its predecessor's 20.3%, and led on EEBench and the Harvey legal-agent benchmark while trailing Fable 5.1 on CursorBench 4.0 and HealthBench Professional.
  • A rebuilt safety stack blocked all but 3.3% of high-risk dual-use prompts in HackerBench v0.3, with few false rejections, and select cybersecurity partners received invitation-only red-team access.
  • Hours before launch, Grok 4.6's three setups lost to every Codex and Claude configuration in Ben Swerdlow's Brood War Bench, managing a best result of 2 wins in 18 matches, and Grok 4.7 has not been ranked.
Elon Musk's SpaceXAI Launches Grok 4.7 With 71% DeepSWE Benchmark Score

Grok 4.7 Built for Multi-Hour Workloads

SpaceXAI, Elon Musk's artificial intelligence unit, released Grok 4.7 on September 21, describing it as the company's most capable model yet for coding and knowledge work. According to the official announcement, the new system runs on a larger base model than Grok 4.6 and went through longer reinforcement-learning cycles on harder task mixes, with a deliberate emphasis on problems that take several hours to complete.

The company framed the release around endurance. Where earlier models answered in minutes, Grok 4.7 is tuned to grind through assignments that would occupy a human engineer or analyst for hours. SpaceXAI said the model is now markedly better at verifying its own output over long runs, handling very long context windows (the amount of text a model can hold in working view at once), and operating the Grok Bot agent framework natively, the kind of multi-step setup where a model plans, calls tools, and executes work on its own rather than answering a single prompt — upgrades it says are visible in conversation, general knowledge tasks, and the generation of professional documents and presentations.

On the published benchmarks, the xhigh configuration sits close to the frontier. On DeepSWE v1.1, a software-engineering suite, Grok 4.7 posted 71.0%, edging past Fable 5.1 max at 70.0% and trailing only GPT-5.6 Sol max at 727%. On Terminal-Bench 4.0, which scores multi-hour terminal work, the model jumped to 38.0% from its predecessor's 20.3%. On suites that simulate multi-hour professional jobs for lawyers, nurses, and financial analysts — AA Briefcase and HealthBench among them — SpaceXAI said Grok 4.7 now runs level with the industry's best.

Safety is the other headline claim. A rebuilt safety stack makes this the company's toughest version yet at refusing improper requests and resisting jailbreaks, the attempts to talk a model past its safeguards. In HackerBench v0.3, which tests malicious cyber tasks, only 3.3% of high-risk dual-use prompts — requests with both legitimate and harmful applications — got through, with few false rejections of legitimate defensive security work. Select cybersecurity partners have been granted invitation-only red-team access, adversarial testing in which security researchers probe a system for weaknesses, to support defense research.

The model is live now through Cursor, Grok Build, the Grok API, major cloud platforms, and model routers, the services that hand a developer's request to whichever AI model they choose through a single interface.

$2 Pricing Held as Brood War Bench Embarrasses Grok 4.6

Pricing is the quiet constant: Grok 4.7 launches at $2 per million input tokens and $6 per million output tokens — identical to Grok 4.6 — plus a fast variant that doubles output speed at double the price. Tokens are the units of text that language models read and generate, and API billing is metered on them: input covers what a developer sends, output what the model writes back. The rates undercut the competition sharply: GPT-5.6 Sol lists at $4/$20 per million tokens and Fable 5.1 at $10/$50, putting Grok's pricing at roughly a third to a half of rival levels with faster delivery.

SpaceXAI's own comparison table shows Grok 4.7 leading on electrical engineering (EEBench: 64.0% versus 39.4% for GPT-5.6 Sol and 56.4% for Fable 5.1) and on the Harvey legal-agent benchmark (19.6% versus 2.5% and 6.7%), while trailing Fable 5.1 on software engineering (CursorBench 4.0: 46.3% versus 51.8%) and on clinical reasoning (HealthBench Professional: 56.7% versus 62.1%). Every figure comes from SpaceXAI's published data.

The corporate backdrop matters: SpaceXAI took its current shape in February, when xAI merged into SpaceX, followed by the June acquisition of Cursor, the AI code editor that doubles as one of Grok 4.7's launch channels — part of Musk's empire that spans Tesla (TSLA) and, per earlier reporting, a push to consolidate AI subscription plans within weeks.

Hours before launch, though, a crowd-sourced benchmark dinged the predecessor. Developer Ben Swerdlow's Brood War Bench, in which AI agents play StarCraft: Brood War, circulated on X on September 20: 19 configurations from Codex, Claude, and Grok fought 171 head-to-head matches, and Codex Astra swept all 18 of its games at maximum reasoning. Grok 4.6's three setups lost to every Codex and Claude configuration, with a best result of 2 wins in 18. Real-time strategy is a demanding test for agents because economy, production, and combat decisions all arrive continuously — there is no turn in which to pause and plan.

Swerdlow logged one match in which a top-reasoning Grok setup burned 11,138 reasoning tokens over 43 minutes yet issued only six command batches — never fielding a single combat unit. He wrote that Grok is not yet smart enough for the game, with older models treating real-time strategy as turn-based, and added that no entrant surpassed novice level. Grok 4.7 has not been ranked.

COINOTAG's Read: Pricing Pressure Is the Real Signal

Taken together, the launch-day numbers and the Brood War results tell one story: frontier labs now sell reasoning at commodity prices, and agentic reliability — not static benchmark scores — decides who wins. In COINOTAG's assessment, the $2/$6 schedule gives Grok 4.7 tokenomics that rivals must match, much the way futures pricing forces markets to re-anchor expectations, while labs without comparable compute risk becoming exit liquidity in this price war. The outlet added that Musk compute headlines have historically moved sentiment in Musk-linked assets such as Dogecoin (DOGE), and advised watching Grok 4.7's agent results rather than its launch claims.