xAI Releases Grok 4.7 With New Safeguards and Higher Multi-Hour Task Performance
Key Takeaways
- •Grok 4.7 is built on a new, larger base model and trained with an extended reinforcement-learning run using harder tasks, which xAI says improves code generation, long-context handling, and self-verification.
- •The model's largest benchmark gains came in long-running coding and terminal work, with Terminal-Bench 4.0 rising to 38.0% from 20.3% and CursorBench 4.0 rising to 46.3% from 40.4%.
- •Grok 4.7 ships with an entirely new safety stack that xAI described as its strongest tested for refusal behavior and jailbreak resistance, and it leads LatchBio's biosafety benchmark at 62.4%.
- •Pricing stays at Grok 4.6 levels of $2 per million input tokens and $6 per million output tokens, below the rates of GPT-5.6 Sol and Fable 5.1.
- •Artificial Analysis scored Grok 4.7 at 46 on its Intelligence Index, a two-point gain that placed SpaceXAI among the top four AI labs, but found the model used roughly 81,000 output tokens per task, more than double Grok 4.6's 36,000, increasing effective costs.

xAI has released Grok 4.7, describing it as its most capable model to date for coding and knowledge work. The company positioned the model as a frontier offering at a competitive price, with pricing unchanged from its predecessor, Grok 4.6.
Grok 4.7 is built on a new, larger base model and was trained through an extended reinforcement-learning run using a more difficult task mix weighted toward problems requiring many hours to complete. According to xAI, the training changes improve code generation, long-context handling, and self-verification. The model also natively understands the Grok Bot harness, which the company said makes it more effective for conversational tasks and general knowledge work.
The largest reported gains were concentrated in longer-running coding and terminal tasks. On CursorBench 4.0, Grok 4.7 scored 46.3%, compared with 40.4% for Grok 4.6. Its score exceeded GPT-5.6 Sol’s 41.7% but remained below Fable 5.1’s 51.8%. At high effort on DeepSWE v1.1, Grok 4.7 scored 71.0%, trailing GPT-5.6 Sol at 72.7% while narrowly exceeding Fable 5.1 at 70.0%.
The model also improved over Grok 4.6 on GDPval and AA Briefcase, which evaluates multi-hour office tasks completed by professionals including lawyers, nurses, and financial analysts. On AA Briefcase, Grok 4.7 scored 1,657, close to Fable 5.1’s 1,678. Its most pronounced benchmark increase came on Terminal-Bench 4.0, which measures multi-hour terminal work. Grok 4.7’s score rose from 20.3% for Grok 4.6 to 38.0%.
On HackerBench v0.3, which covers risky cyber tasks, the model allowed 3.3% of dangerous dual-use prompts through while rarely blocking legitimate security work.
Grok 4.7 is here. It's a notable improvement over Grok 4.6 at the same price and speed. pic.twitter.com/H3OTBbXyvO — SpaceXAI (@SpaceXAI) September 21, 2026
https://x.com/SpaceXAI/status/2102069815225586149?ref_src=twsrc%5Etfw
New safety stack and unchanged pricing
A central feature of the release is Grok 4.7’s safety stack, which xAI described as entirely new and the strongest it has tested for refusal behavior and resistance to jailbreaks. The model leads LatchBio’s biosafety benchmark with a score of 62.4%, combining performance on benign tasks with refusals on dangerous requests in dual-use fields such as cybersecurity and biological research. xAI has also started providing select cybersecurity partners with invite-only access to the model’s red-team capabilities for defense research.
Grok 4.7 is available at the same rates as Grok 4.6: $2 per million input tokens and $6 per million output tokens. Those rates are below GPT-5.6 Sol’s $4 and $20 prices, respectively, and Fable 5.1’s $10 and $50 prices. A fast variant with twice the output speed is available at twice the price. Flat per-token rates make the sticker comparison straightforward, though the total cost of long-running work also depends on how many tokens a model consumes to finish a task.
On CursorBench’s price-performance frontier, the pricing places Grok 4.7 among the more cost-efficient options for extended coding workloads. The model is available immediately in Cursor and Grok Build, as well as through the Grok API, third-party coding harnesses, model routers, and cloud platforms.
The release adds to competition among frontier-model providers, which are increasingly comparing systems on long-duration agentic tasks, safety calibration, and cost per completed task in addition to headline benchmark scores. For developers and enterprises assessing coding-oriented models, Grok 4.7 combines stable pricing with higher multi-hour task scores and new safeguards.
Artificial Analysis confirms gains but flags token consumption
Independent benchmarking firm Artificial Analysis also evaluated Grok 4.7, giving it a score of 46 on the Intelligence Index. That was two points higher than Grok 4.6 and placed SpaceXAI among the top four AI labs, according to the firm. The outside assessment lines up with xAI’s positioning of the release as a step forward on coding and knowledge work.
The evaluation used xhigh reasoning effort and identified the clearest progress in long-horizon agentic knowledge work. On Artificial Analysis’ private AA-Briefcase benchmark, Grok 4.7 gained 111 Elo points over its predecessor to reach 1,657, placing it alongside Claude Opus 5 and Claude Fable 5.1 at the frontier. The increase was driven primarily by analytical quality, which rose to 1,994 Elo from 1,690. Presentation quality remained roughly stable.
On GDPval-AA, which measures practical work products such as documents, spreadsheets, and slides, Grok 4.7 scored 1,695 Elo, an increase of 90 points.
Grok 4.7 scores 46 on the Artificial Analysis Intelligence Index to bring SpaceXAI into the top 4 AI labs. Coding Agent Index performance has also improved, overtaking GPT-5.6 Sol Grok 4.7 scores +2 points over Grok 4.6 on the Intelligence Index, with strong performance on… pic.twitter.com/8p4KC6cgZy — Artificial Analysis (@ArtificialAnlys) September 21, 2026
https://x.com/ArtificialAnlys/status/2102074898327932987?ref_src=twsrc%5Etfw
Artificial Analysis found that the performance improvements required substantially more compute. At xhigh effort, Grok 4.7 used approximately 81,000 output tokens per Intelligence Index task, more than twice the 36,000 tokens used by Grok 4.6 and nearly three times the 27,000 tokens consumed by GPT-6 Astra at maximum effort. Because per-token rates are unchanged, that heavier drawdown raises the effective cost of a comparable task even though headline prices stayed flat.
The firm reported modest additional improvements on Terminal-Bench 4.0 and GDP.pdf, along with small regressions on AA-LCR and AutomationBench-AA. Reliability improved modestly: the AA-Omniscience hallucination rate fell to 29% from 34%, while accuracy remained nearly unchanged at 47%.
Other technical specifications remain the same as in the previous generation, including a 500,000-token context window. Cache-hit pricing is discounted to $0.50 per million tokens. With the model live across Cursor, Grok Build the Grok API, third-party harnesses, model routers, and cloud platforms, developers can now weigh the benchmark gains against the higher token consumption on their own workloads. The original report is available from Metaverse Post.