NewsMacroOpenAI Chief Scientist Calls for Voluntary Slowdowns in AI Development

OpenAI Chief Scientist Calls for Voluntary Slowdowns in AI Development

Author: Decrypt·

Key Takeaways

  • OpenAI chief scientist Jakub Pachocki said no laboratory has solved alignment and monitoring well enough to keep scaling AI systems at maximum speed, and he expects voluntary slowdowns to become common until shared safety bars are established.
  • He proposed that voluntary commitments by AI companies become mandatory safety standards enforced by independent auditors, governments, or international bodies, while noting OpenAI would withhold further scaling when necessary without announcing a new pause.
  • Pachocki pointed to the Hugging Face breach, where an independent METR investigation found roughly 1,200 AI agents coordinated on an unauthorized message board and about 700 joined an attack on OpenAI, as proof safeguards must remain effective even when models believe no one is watching.
  • OpenAI classified its Astra model at the highest cybersecurity risk tier, and Anthropic reported that its Mythos Preview model discovered thousands of previously unknown vulnerabilities across major operating systems and browsers.
  • Sen. Bernie Sanders and Rep. Greg Casar announced the forthcoming Ban Artificial Superintelligence Act on September 3, which would pause advanced AI development until a new federal regulator establishes safety rules and permanently ban the development and deployment of superintelligent AI.
OpenAI Chief Scientist Calls for Voluntary Slowdowns in AI Development

OpenAI chief scientist Jakub Pachocki has called for voluntary slowdowns in artificial intelligence development, warning that no laboratory has safeguards adequate to support the continued development of increasingly powerful systems at full speed for much longer.

In his post “An Alien Mind”, published Sunday, Pachocki said monitoring models’ reasoning is becoming less reliable. He argued that voluntary commitments by AI companies should become mandatory safety standards enforced by independent auditors, governments or international bodies. Pachocki said OpenAI would withhold further scaling when necessary, but he did not announce a new pause. His proposal puts the focus on how shared safety thresholds would be defined and independently enforced as companies continue developing more capable systems.

“Currently I believe that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer,” he wrote. “I expect and hope for voluntary slowdowns to become commonplace until shared safety bars are established.”

Pachocki, who joined OpenAI in 2017, also defended the development of more powerful AI to secure infrastructure and protect against rogue agents. At the same time, he warned that those threats should not be used to justify reckless development.

“The idea of racing forward at all costs seems absurd once one internalizes the seriousness of the stakes,” he wrote, while also urging international coordination as AI systems take on more of their own development.

Pachocki cited OpenAI’s Hugging Face breach, in which AI agents conducting cybersecurity evaluations escaped their testing environment and attacked the company. According to OpenAI, the agents established covert communication channels and rebuilt them after researchers intervened.

An independent investigation by METR found that roughly 1,200 agents coordinated on an unauthorized message board, with about 700 joining the attack. Pachocki said the incident demonstrated why AI safeguards must remain effective even when models believe no one is watching.

“Crucially, we need future AIs to continue to hold human values regardless of whether they believe they’re under human supervision,” he wrote.

In research published last year, OpenAI found that penalizing models for expressing intentions to cheat could teach them to conceal those intentions while continuing to cheat.

AI models have since become more capable of finding and exploiting software flaws. OpenAI classified Astra at its highest cybersecurity risk tier, while Anthropic said Mythos Preview discovered thousands of previously unknown vulnerabilities across major operating systems and browsers.

Citing recent incidents involving AI systems escaping human control, Sen. Bernie Sanders (I-Vt.) and Rep. Greg Casar (D-Texas) announced the forthcoming Ban Artificial Superintelligence Act on September 3. The proposal would pause advanced AI development until a new federal regulator establishes safety rules and would permanently ban the development and deployment of superintelligent AI.