Image: Reflection AI
Reflection AI has unveiled Beam, its first open-weight model and its first big step toward giving the US a homegrown rival to China's open AI models. Beam is a Mixture-of-Experts model with 501 billion total parameters, but only 23 billion are active for each token.
Introducing Beam: a highly efficient agentic open model with 501B total parameters and 23B active.
— Reflection (@reflection_ai) October 5, 2026
- Frontier reasoning efficiency
- Advances the Western open frontier on coding & agentic tasks
- Trained end-to-end from scratch
Full weights release this month.
Learn more about… pic.twitter.com/1rMABCywUG
You can't download it yet. Reflection says Beam is going through final red-teaming and testing, with early access for a small group via a waitlist. The weights, a technical report and a model card are due later this month under the Apache 2.0 license, which allows commercial use.
Reflection was founded in 2024 by former Google DeepMind researchers Misha Laskin and Ioannis Antonoglou. It started out building coding agents, then repositioned itself as America's open frontier AI lab and a Western answer to DeepSeek. It raised $2 billion at an $8 billion valuation from backers including Nvidia, Sequoia and Lightspeed, and Laskin later confirmed a new round at a $25 billion pre-money valuation.
Beam is text-only and built for coding, reasoning and agent work. Reflection says it trained the model from scratch on 23.8 trillion tokens and extended its context window to 1 million tokens.
It then ran what it calls one of the largest reinforcement learning runs by any open lab: more than 100 million practice attempts on 10,500 Nvidia GB300 GPUs over four weeks.
The 23B figure is the one to watch. In a Mixture-of-Experts model, each token passes through only a small slice of the network, so the compute per token tracks active parameters, not the total.
Reflection says Beam scores about the same as GLM-5.2 on advanced reasoning tests while using three to four times less inference compute. That's an estimate based on active parameters and tokens generated, not measured serving costs.
There's a catch for anyone planning to run it themselves. All 501B parameters still have to sit in memory. At standard 16-bit precision, that's roughly a terabyte for the weights alone, so Beam is cheap per answer but still needs a serious multi-GPU server.
Reflection's numbers back up its "Western open frontier" claim. Beam scores 80.9 on SWE-bench Verified, ahead of Thinking Machines' Inkling (77.6) and Nvidia's Nemotron 3 Ultra (70.7). On Terminal Bench 2.1, it posts 80.1 against 63.8 and 56.4. Both rivals use far more active parameters: 41B for Inkling and 55B for Nemotron 3 Ultra.
Image: Reflection AI
But "Western open frontier" is not the same thing as the open frontier. Reflection's full results table shows Kimi K3 (88.3), GLM 5.3 (88.2) and DeepSeek V4.1 Flash (90.6) all ahead of Beam on Terminal Bench 2.1. Reflection admits Kimi K3 stays ahead on raw capability.
None of those three appear in the chart Reflection posted on X. OpenAI's gpt-oss and Meta's Llama aren't in either comparison.
So why does a US-built open model matter if Chinese ones score higher? Reflection's business depends on that answer. Laskin has said revenue will come from large enterprises and from governments building "sovereign AI," and that big companies want a model "you will have ownership over" that runs on their own infrastructure. He also told Semafor he wants open-model builders to have a voice in Washington's AI regulation talks.
Until the weights ship, every Beam number is Reflection's own. The company took rival scores from trackers Artificial Analysis and DataCurve, and some differ from the makers' own figures. Nvidia lists Nemotron 3 Ultra at 71.9 on SWE-bench Verified, for example. Reflection's X chart also shows Beam at 15.6 on CritPt, while its blog table says 16.3.
Next up: the technical report, promised safety evaluations and independent tests. Reflection says it is already training its next model.


Comments