Reflection AI Opens Early Access to Beam Ahead of October Open-Weights Release
Key Takeaways
- •Reflection AI launched early access to its first open-weight model, Beam, on October 5, with the weights, technical report, and model card to follow later this month under an Apache 2.0 license that permits commercial use and modification.
- •Beam is a sparse mixture-of-experts model with 501 billion total parameters and 23 billion active per token, and Reflection says it matches Z.ai GLM-5.2 on advanced reasoning tests while using 3-4 times less inference compute, though these claims remain unverified.
- •Beam's reinforcement learning training run used 10,500 Nvidia GB300 GPUs over four weeks, generated more than 100 million rollouts, and employed roughly 1.3 billion sandboxes, which Reflection describes as one of the largest RL runs by any open lab.
- •On Reflection's benchmark tables, Beam outperformed Thinking Machines Lab's Inkling on four coding benchmarks, scoring 77.2 versus Inkling's 56.9 on SWE Bench Pro v2-Hard.
- •Founded in 2024 by former Google DeepMind researchers Misha Laskin and Ioannis Antonoglou, Reflection has raised about $4.7 billion at a $25 billion pre-money valuation, and the open-weight release will make independent third-party evaluations of Beam possible.

Reflection AI opened early access to Beam, its first open-weight model, on October 5, according to the company's official announcement. The model is currently completing red-team testing — a pre-release process in which adversarial testers probe a model for safety and security flaws — and the weights, a technical report, and a model card are scheduled to follow later this month.
According to Reflection, the weights will ship under an Apache 2.0 license, together with the full stack required to run, evaluate, and fine-tune the model. Apache 2.0 is a permissive license that permits commercial use, modification, and redistribution, and the open-weights release will let developers download the parameters and run the model on their own hardware.
GLM-5.2 parity with 3–4× less inference compute
Beam is a sparse mixture-of-experts model with 501 billion total parameters and 23 billion active parameters per token. In such architectures, only a subset of the model's parameters is activated for each token, reducing the compute required per token relative to the total parameter count. It was pre-trained on 23.8 trillion tokens, and its context length reached 1 million tokens during mid-training.
Reflection says Beam matches Z.ai's GLM-5.2 on advanced reasoning tests while using “3–4× less inference compute.” The company's claims have not been independently verified, though the open release will make third-party evaluations possible. GLM-5.2 has approximately 744 billion parameters, 40 billion of them active. Reflection places Kimi K3 ahead of Beam in raw capability and pitches Beam's advantage as inference efficiency.
Against Thinking Machines Lab's Inkling, which was released in July, Reflection's benchmark tables show Beam ahead on four coding benchmarks, with scores published for both models. On SWE Bench Pro v2-Hard, a benchmark that measures real-world software engineering performance, Beam scored 77.2 to Inkling's 56.9. Inkling is multimodal, while Beam is text-only.
10,500 Nvidia GB300 GPUs ran for four weeks of RL
Reflection's reinforcement learning run used 10,500 Nvidia GB300 GPUs over four weeks, produced more than 100 million rollouts, and utilized about 1.3 billion sandboxes for training and grading. Rollouts are complete model attempts at a task, executed and graded inside isolated sandbox environments. The company describes it as one of the largest RL runs by any open lab. By Reflection's count, Inkling trained on 30 million rollouts.
Reflection was founded in 2024 by Misha Laskin and Ioannis Antonoglou, both former Google DeepMind researchers. The company has raised about $4.7 billion from investors including Nvidia, Sequoia Capital, and Lightspeed Venture Partners, most recently at a $25 billion pre-money valuation.
A year ago, it raised $2 billion at an $8 billion valuation, as Cryptopolitan reported. Its SpaceX deal alone runs up to $6.3 billion, $150 million per month for Nvidia GB300 capacity at Colossus 2 near Memphis, Cryptopolitan reported in June.
Beam also underpins Reflection's “AI factory” offering, which lets institutions train its models on their own data and run them on their own compute Shinsegae Group is trialing a sovereign edition in South Korea.
“They're kind of like rocket ships,” Laskin has said of his models' path to the frontier. “To build a big rocket ship, it takes time.”