OpenAI Quietly Revised Multiple GPT-6 Astra Benchmark Scores After Launch, With Some Figures Still Changing
Key Takeaways
- •OpenAI changed several GPT-6 Astra benchmark figures after publishing and republishing its Sept. 3 announcement blog post, with some revisions showing Astra performing better and rival Anthropic models appearing worse.
- •Astra's reported hallucination rate was briefly cut from 4.2% to 2% before returning to the original figure, and Astra's ARC-AGI-3 score rose from 98.6% in an embargoed draft to 99.99% in the live blog.
- •OpenAI's ARC-AGI-3 result depended heavily on configuration: the Arc Prize Foundation independently measured 99.9% with a powerful harness versus 63% with the standard harness.
- •OpenAI said evaluation numbers carry a few percentage points of noise and that it made fixes to represent its best estimate of model performance, while some experts raised concerns about industry-wide 'benchmaxxing' practices.
- •Changing and confusing benchmark figures could make it harder for customers and investors to assess which AI models excel at which tasks, particularly as OpenAI eyes a possible 2027 IPO.

OpenAI has altered several evaluation benchmarks for its GPT-6 Astra model since first publishing the model's announcement blog post in the mid-afternoon of Sept. 3. In some instances, the updated figures showed Astra performing better, while scores for models from OpenAI's chief rival Anthropic appeared worse.
The changes came during an unusually rocky rollout of the blog post. OpenAI had originally planned for the post to go live at 2 p.m. ET, but it took nearly two more hours before it was widely viewable online.
When OpenAI's X account shared the blog post at 3:32 p.m., the link failed to load properly, returning an error message. At 3:50 p.m., OpenAI CEO Sam Altman posted the link, writing, "We hit a little snag getting the blog post deployed, but it is really great." Multiple commenters were still unable to view the page and received the same error, as did Fortune. When Fortune checked back roughly an hour later, the post was visible and loading properly.
It later emerged that OpenAI had actually published the blog shortly after 2 p.m. but then retracted it for reasons the company said it could not disclose, while stating those reasons were unrelated to the benchmark performance figures. (OpenAI first told Fortune it was a bug in the content management system, and later attributed it to an internet outage.) Upon republishing, the blog contained different evaluation metrics that appeared to favor Astra—and some figures have continued to change since. The revisions are documented through internet archive snapshots, which capture pages at specific timestamps and make it possible to compare earlier versions of a post against later ones.
The disclosure of these changes comes amid fierce competition in the AI industry, where companies release updates to their large language models at a frenetic pace, each trying to pull ahead of rivals. The focus on metrics also underscores the difficulty of measuring large language model performance with standardized benchmark tests, along with concerns that such specifications are prone to manipulation and gamesmanship.
"We care deeply about getting evaluations right," an OpenAI spokesperson told Fortune. "Most evaluations have noise within a few percentage points based on the exact checkpoint, scaffold, and evaluation run used in reporting. For our launch blog, we made fixes to ensure the numbers represent our best estimate of available model performance, so that users can make meaningful comparisons."
Discrepancies Between the First and Final Published Blogs—and the Numbers Are Still Changing
One of the most notable changes involved Astra's reported hallucination rate. In the first internet archive snapshot of the blog post, from 2:23 p.m. ET, it was 4.2%. It stayed at that number through several more snapshots, the last being a fifth at 3:11 p.m. ET—about 10 minutes before OpenAI tweeted out the final version.
But the hallucination rate, along with four other metrics, changed in the sixth archival snapshot, taken at 5:20 p.m.—after the blog was presumably visible to everyone. The rate was halved to 2% for Astra. The score for Astra's predecessor, GPT-5.6 Sol, also fell from 12.2% to 9.4%. OpenAI has continued to change this metric; as of this writing, the hallucination rates are back up to their original 4.2% and 12.2%.
OpenAI also appears to have given GPT-5.6 Sol a large boost on its internal version of the ExploitBench cybersecurity evaluation, rising from 5.5% in the first version to 11.5% in later versions. OpenAI said it is currently investigating reverting that figure to 5.5%, because it says the 11.5% result reflects a reasoning level that is not commercially available for Sol.
Astra is especially strong at mathematics, OpenAI says—a quality the company highlights in the opening paragraph of the announcement page. While that metric did not change across the snapshots for Astra—it remains 97.6% on the FrontierMath Tier 4 (v2) eval—OpenAI did briefly alter the scores for GPT-5.6 Sol and Anthropic's latest model, Fable 5.1. Those changes briefly made Astra appear significantly better at math than both models.
In the first snapshot (2:23 p.m. on Sept. 3), Anthropic's Fable 5.1 scored 87.8%. By 5:17 p.m., it had dropped nearly 10 percentage points to 78%. Today, it is back up to 83%. Similarly, GPT-5.6 Sol's score went from 83%, down to 80.5%, and back up to 83% today.
The metric changes began even before OpenAI first published the blog at 2 p.m. An embargoed pre-publication draft the company provided to Fortune and other media organizations listed Astra's score on the ARC-AGI-3 evaluation as 98.6%. In the live blog, it is now 99.99%.
"We always verify evals before publication so adjustments between draft and final version are normal," a company spokesperson said at the time. OpenAI also noted that the benchmark's creator, the Arc Prize Foundation, found that Astra performed at 99.9% in its independent assessment, provided the model was given a particularly powerful harness (a set of tools the model can use to complete tasks). With the benchmark's standard harness, it performed at 63%—still significantly better than any other AI model currently in public release. OpenAI said "things like harness, reasoning level and other factors inform evals." That gap between the two ARC-AGI-3 figures illustrates how heavily a headline number can depend on the test configuration used to produce it.
"Benchmaxxing"—or Improving Accuracy?
Different research teams at OpenAI oversee different metrics and are responsible for calculating and reporting them to a central team for publication. OpenAI is open about the fact that the numbers are achieved under the best possible conditions and may differ slightly from the models available in the production ChatGPT product that most users access. "Evaluation scores are the maximum at any effort," reads a disclaimer on the blog. The company includes further caveats on each metric in footnotes.
Accuracy is elusive, as multiple numbers can be considered accurate depending on the conditions under which the tests were run. But some AI experts question whether "benchmaxxing" is also at play. This is a known practice in the AI industry—not just at OpenAI—of maximizing scores by re-running evaluations under different conditions.
"This can be done in a very tight timeframe, and it's better for their marketing," said Anka Reuel and Mike Hardy, researchers at the Stanford Intelligent Systems Laboratory and Stanford Trustworthy AI Lab. They also noted that the GPT-6 Astra system card, which should contain more technical information on how the evaluations were performed, does not always properly explain them. For the internal hallucination benchmark, for example, the system card provides "barely any details about the evaluation," they said. "It doesn't even include the number of test items."
This re-running of numbers could explain why Astra's coding capabilities also received a marginal boost in later versions of the blog post, rising from 57.7% to 57.9%. Though the difference is negligible, OpenAI appeared to care enough about it to swap in the new, improved figure.
Not every change OpenAI made portrayed Astra more favorably. For example, two Anthropic model scores improved across different versions of the healthcare-focused eval HealthBench Professional: Claude Fable 5.1 rose from 56.6% to 58.1%, and Opus 5 went from 54.5% to 56.4%. Scores for models made by other AI companies are typically taken from published leaderboards and do not involve OpenAI running assessments on rivals' models itself.
Evaluation Score Debates Haunt the AI Industry
The question of benchmark accuracy has surfaced repeatedly. In 2025, Meta denied reports that it artificially boosted scores for its Llama 4 model by publishing results from an internal version of the model rather than the publicly available one. Yann LeCun, Meta's former chief AI scientist, later admitted the company had "fudged" the benchmark results. Evaluation metrics also change frequently as new ones are created. ExploitGym, a cybersecurity benchmark at the center of the July incident in which OpenAI's models went rogue and attacked the company Hugging Face, was created in 2026.
Vincent Sunn Chen, an AI engineer who leads benchmark and evaluation research at Snorkel AI, said it is not unusual for benchmark scores to shift in the final hours before a model launches. "It's usually a function of final launch logistics," he said in an email. "A benchmark score reflects a specific measurement setup: the model checkpoint, configuration (including how much time and compute the model is allowed), harness, eval/grading configuration (e.g., non-determinism in the judge). All of those are typically still shifting in the final days before a launch, so I'm not surprised that there were some updates." He said he would like to see industry norms develop requiring companies to report what has changed about an assessment when they revise benchmark performance numbers, so researchers can interpret the results more clearly.
Benchmark results matter for several reasons. They are how AI companies measure progress—but also a way of keeping score in the race against competitors. Topping leaderboards can help AI companies win customers and, in some cases, help them hire engineers and researchers. Yet as this episode illustrates, interpreting benchmark scores can be technically complex, posing a challenge for companies that want to present results to the public in a digestible format. These complexities, combined with confusion over changing metrics and accusations that companies have not been intellectually honest in how they present results, could make it harder for customers and investors to determine exactly which models are best for which tasks. The confusion could also muddy the narrative of having the best models on the market—one that OpenAI would no doubt like to present ahead of a possible 2027 IPO.
This story was originally featured on Fortune.com.