GPT-6 Astra Users Say OpenAI’s New Model Has Been ‘Nerfed’
Key Takeaways
- •Some users report GPT-6 Astra now delivers faster responses with lower-quality output, and developer Pankaj Kumar suspects OpenAI reduced the model's computing effort, a change the company has not confirmed.
- •Salio and researcher Md Ismail Sojal said they ran identical prompts on launch day and the current version with the same settings, and both reported worse results from the current model.
- •Dax Raad's Opencode team switched back to GPT-5.6 Sol after its spending doubled, concluding Astra's drawbacks were not worth the cost of its premium pricing.
- •Critics including Antikythera and T3Chat founder Theo argue the model has not changed and that fading launch-week hype is exposing long-standing inconsistencies.
- •Astra remains OpenAI's first model to cross the company's critical cybersecurity risk threshold and is priced at $10 per million input tokens and $50 per million output tokens, 2.5 times what Sol charged at launch.

Users are flooding X with complaints that OpenAI’s GPT-6 Astra has become less capable only a week after its launch. Some users report faster responses and poorer results, while others argue that the model was always inconsistent and that the initial excitement simply obscured its weaknesses.
A week ago, GPT-6 Astra was rebuilding Manhattan street by street inside a game engine, impressing users with its capabilities. This week, some of the same users are posting screenshots and asking OpenAI what happened to the model.
“Astra feels significantly dumber for me today,” the pseudonymous developer synthwavedd wrote on X. “Was only a matter of time before The Post-Launch Lobotomy. Shame.”
The pattern is familiar to users of artificial-intelligence chatbots and agents, who sometimes feel that a favored model has been “nerfed” after release—a term borrowed from gaming that describes developers weakening something that was overpowered. There can be reasonable explanations for perceived performance declines, but the number of similar complaints about GPT-6 Astra has prompted users to compare outputs and examine whether the model has changed. The question carries practical weight for developers who products on models they rent rather than run, and because proprietary systems expose no internals to inspect, side-by-side output comparisons are among the few verification tools outsiders have.
Developer Pranjal Paliwal, who had previously praised Astra, said he later reviewed code written by the model and changed his assessment.
“We don’t have AGI,” Paliwal wrote. “We have a regression.”
AGI, or artificial general intelligence, is the industry’s term for a machine capable of performing essentially any cognitive task a human can perform. OpenAI’s president used the term at Astra’s launch. A week later, Paliwal invoked it to describe the gap between the model’s apparent capabilities and its mistakes.
Other complaints follow a similar pattern. Developer Pankaj Kumar listed faster answers and lower quality among the symptoms, along with a suspicion that OpenAI had “reduced the juice value.” Founder Saba also asked OpenAI why she now has to “dumb it down” to complete tasks.
“Juice value” is not an official technical term. In this context, users use it as shorthand for the amount of computing effort a model applies before answering, and for the suspicion that OpenAI may have reduced that effort after launch demonstrations. OpenAI has not confirmed that it made such a change to Astra.
Some users have attempted direct comparisons. Salio and researcher Md Ismail Sojal said they ran identical prompts against Astra on launch day and again on the current version, using the same settings. Both reported worse results from the current model.
“Today’s GPT-6 Astra output looks worse. GPT-6 Astra at launch vs GPT-6 Astra today,” Sojal wrote. The post said the comparison used the same prompt and settings and that “[t]he difference is bigger than I expected.”
Other users say they have switched back to Astra’s predecessor. Dax Raad, who builds the coding tool Opencode, said his team returned to GPT-5.6 Sol because its spending had doubled while the disadvantages of using Astra were not worth the cost. ChatGPT user Mustafa Sahinli compared Astra to Claude Opus 4.6 “after 1 week of release,” referring to backlash that Anthropic’s model also faced after launch.
Not everyone believes OpenAI deliberately weakened Astra. The pseudonymous user Antikythera offered a detailed rebuttal, arguing that the timeline may run in the opposite direction.
“It is as dumb as it was on launch,” Antikythera wrote. “The model is good, but the model has a lot of problems. It’s lazy. Writes like a bullet-point-addict... people were overhyped on launch week, now they had time to test it and see its mistakes.”
Under that explanation, Astra did not become less capable; users simply stopped being dazzled by its early demonstrations and began noticing its recurring problems.
Theo, founder of T3Chat, offered a similar view. He said Astra is more inconsistent than Claude Fable and can produce either outstanding or extremely poor code.
“GPT-6 Astra has done incredible things I never thought a model could do. It has also done some of the stupidest things I’ve ever seen a model do. Generally speaking, Fable 5.1 just does what I ask,” Theo wrote on X on September 8, 2026. The increase in negative examples, he suggested, may reflect the end of the model’s initial honeymoon period rather than a confirmed change in its behavior.
The debate echoes events involving OpenAI’s previous flagship model, GPT-5.6 Sol. In July, users said Sol’s highest reasoning mode had suddenly become less thorough. OpenAI executive Tibo Sottiaux denied deliberately weakening the model but confirmed that the company had been experimenting with “reasoning effort,” a setting that controls how many steps a model works through before producing an answer.
One reply summarized the recurring joke about closed AI laboratories: a company releases a model, and a few days later it catches “some kind of disease a few days later and suddenly become[s] dumber.” Another user offered a more cynical explanation: “They probably get quantized so they’re not burning the companies as much money.”
Quantization reduces the precision of a model’s internal calculations to lower computing costs, often with a potential trade-off in accuracy. OpenAI has never confirmed intentionally quantizing a shipped model.
OpenAI has not issued a Sol-style statement about Astra. In July, a similar wave of Sol complaints drew an on-record response from a company executive; whether OpenAI addresses the Astra complaints the same way remains an open question. The model remains the company’s first to cross what OpenAI calls the critical threshold for cybersecurity risk. That designation means Astra can find and chain together previously unknown software vulnerabilities without human guidance, a capability restricted to vetted defenders through OpenAI’s Daybreak program.
Regardless of whether Astra has become less capable, its listed price remains $10 per million input tokens and $50 per million output tokens—2.5 times what Sol charged at launch, a premium at least one team has already concluded is not worth paying.