Reflection opens 501B Beam model for coding and agents
Reflection AI has released Beam, a 501-billion-parameter open-weight model that targets coding, reasoning and long-horizon agent workloads.
A large open-weight bet
Reflection AI has introduced Beam, a sparse mixture-of-experts model with 501 billion total parameters and 23 billion active parameters. The company positions its first open-weight release around coding, reasoning and agentic workloads rather than general chat alone.
Reflection says Beam was pretrained on 23.8 trillion tokens drawn from curated web material and licensed datasets. Its reinforcement-learning program generated more than 100 million rollouts on 10,500 NVIDIA GB300 GPUs over four weeks. Those figures place the model among the most compute-intensive open releases yet publicly described.
The architecture matters as much as the headline parameter count. Activating 23 billion parameters per token gives Beam a chance to offer large-model capacity without paying the full inference cost of a dense 501-billion-parameter system. That design is increasingly important as coding agents consume long contexts, call tools repeatedly and require sustained performance rather than a single benchmark answer.
Reflection says the model is still undergoing final red-teaming and evaluation. The company has therefore released an open-weight artifact before the market has a complete, independently verified picture of its safety, reproducibility or real-world coding performance. Early comparisons reported by third parties are useful signals, but they do not settle whether Beam can compete with established open models across diverse production workloads.
Why it matters
Beam raises the ceiling for open-weight agent models while exposing the cost of reaching that ceiling: more than ten thousand GB300 GPUs and a large-scale reinforcement-learning pipeline. If the weights prove usable outside Reflection’s own stack, the release could give developers another serious alternative to proprietary coding models and intensify competition around efficient MoE inference.