xAI's Grok 4.7 Outscores Anthropic and OpenAI Models on Legal Agent Benchmark at a Fraction of the Price
Key Takeaways
- •Grok 4.7 scored 19.6% on the Harvey Legal Agent Benchmark, roughly triple Anthropic's Fable 5.1 at 6.7% and nearly eight times OpenAI's GPT-5.6 Sol 2.5%.
- •The model is priced at $2 per million input tokens and $6 per million output tokens, making its input cost one-fifth of Fable 5.1's $10 and half of GPT-5.6 Sol's $4.
- •Anthropic's Fable 5.1 keeps substantial leads on longer coding and terminal benchmarks, winning Terminal-Bench 4.0 by roughly 20 percentage points over Grok 4.7.
- •Grok 4.7 runs on 2.1 trillion parameters, 40% more than its predecessor, and was trained with supplemental SpaceX data including Starlink satellite telemetry and manufacturing records.
- •All published benchmark figures come from xAI's own testing rather than independent verification, and the CursorBench behind its price-performance chart is made by Cursor, which SpaceX recently acquired.

xAI launched Grok 4.7 on Monday, introducing a new flagship model that outscored Anthropic's Fable 5.1 and OpenAI's GPT-5.6 Sol on a legal-agent benchmark while charging a fifth of Fable 5.1's price per input token.
According to xAI's release announcement, Grok 4.7 scored 19.6% on the Harvey Legal Agent Benchmark, where Fable 5.1 reached 6.7% and GPT-5.6 Sol managed 2.5%. That puts Grok 4.7 at roughly three times Anthropic's score and close to eight times Sol's. Cryptopolitan reported last month that the company's previous model, Grok 4.6, had already led the same pair at 15.8%. The Harvey benchmark is maintained by legal-technology company Harvey and scores models on multi-step legal work, one of the professional domains where agentic AI adoption is watched most closely.
Legal work stands out as one of the few domains in which Grok 4.7 outperforms both competitors outright. The model also clears Fable 5.1 on the EEBench electrical engineering test and edges past it on the DeepSWE coding benchmark.
Pricing and price-performance
Grok 4.7 is priced at $2 per million input tokens and $6 per million output tokens — the same rates as Grok 4.6. That input price is half of GPT-5.6 Sol's $4 and one-fifth of Fable 5.1's $10. On the output side, Grok 4.7 charges $6, versus $20 for Sol and $50 for Fable, or about an eighth of Anthropic's rate. Per-million-token rates are the standard way API providers bill — input tokens cover what a model reads, output tokens what it writes — so these list prices are the figures developers weigh when choosing between frontier models.
xAI also plotted CursorBench 4.0 scores against the average cost per completed task and claims the model sits on the price-performance frontier. Grok 4.7 reaches around 46% on that chart at roughly $6 per task, while Claude Opus 5 needs almost double the spend to achieve a similar outcome. Fable 5.1 performs better at higher budgets, climbing to 51.8% at around $17 per task.
Elon Musk described the release as "a strong combination of intelligence, speed & low cost" in a post on X.
Anthropic retains the lead on terminal and coding tests
Grok 4.7 finished runner-up to Fable 5.1 on GDPval and on the AA Briefcase office-work test, and Anthropic's model keeps a clear lead on longer coding and terminal benchmarks. The widest margin comes on Terminal-Bench 4.0, where Fable 5.1 scored 57.9% against Grok 4.7's 38.0% — a gap of roughly 20 points. Fable also leads on CursorBench and on HealthBench Professional clinical reasoning, where GPT-5.6 Sol likewise beats Grok.
Grok 4.7 earned 1,695 Elo on GDPval, a benchmark that measures models on tasks performed by lawyers, nurses and financial analysts, up from Grok 4.6's 1,605. That trails Fable 5.1's 1,735 but sits ahead of the 1,542 recorded by OpenAI's newer GPT-6 Astra on the same chart. Elo ratings, adapted from chess, rank models on relative results rather than raw percentage scores.
Taken together, Grok 4.7 is ahead of GPT-5.6 Sol on five of the seven benchmarks, losing only on DeepSWE and the clinical test.
A larger model trained with SpaceX data
Grok 4.7 runs on 2.1 trillion parameters, 40% more than the 1.5 trillion powering Grok 4.6. xAI has added supplemental SpaceX training data, including Starlink satellite telemetry and manufacturing records, and says the new model is also more likely than its predecessor to spend extra time on arduous problems and to double-check its own answers.
The model launched on the Grok app, Cursor, Grok Build and the xAI, with no waitlist.
All figures above come from xAI's own testing rather than independent verification. CursorBench, the benchmark behind xAI's price-performance chart, is made by Cursor, which SpaceX finished acquiring last month. xAI benchmarked Grok 4.7 against GPT-5.6 Sol, while OpenAI's newer GPT-6 Astra appears only in the GDPval, AA Briefcase and EEBench comparisons. Third-party evaluations will be the first outside check on those margins.