TL;DR
Get tech for your team delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Reflection has introduced Beam, its first open-weight model, with 501 billion total parameters and 23 billion active per token. The company reports strong coding and agentic benchmark results and lower inference compute on some reasoning comparisons; the model is still undergoing red-teaming, and its weights are not yet released.
Reflection has introduced Beam, its first open-weight model, a sparse Mixture-of-Experts system with 501 billion total parameters and 23 billion active per token. The company says it is designed for coding, reasoning and agentic workloads; the model is still undergoing final red-teaming and evaluations, and its weights have not yet been released.
Reflection says Beam was pretrained on 23.8 trillion curated tokens drawn from the web and proprietary licensed datasets. The company describes its training as a combination of large-scale pretraining and reinforcement learning, or RL. It says the model matches or outperforms available open base models of similar size, though the announcement does not provide a complete independent evaluation of that comparison.
For its RL campaign, Reflection reports using 10,500 NVIDIA GB300 GPUs for four weeks, generating more than 100 million rollouts. The company also says training and grading used approximately 1.3 billion sandboxes and that it sourced one million coding, agentic and STEM environments. These figures describe Reflection’s reported training process; the release does not provide an external audit of them.
On coding and agentic benchmarks, Reflection says Beam is competitive with larger open models, including GLM 5.2, and approaches Qwen 3.8-Max on some tasks. Its published table lists scores such as 80.1 on Terminal Bench 2.1 and 80.9 on SWE-bench Verified. The results shown vary by benchmark and competitor, and several entries are marked as unreported. Reflection says it will publish the weights, technical report, model card and developer materials later this month.
Lower Compute for Coding Tasks
Beam’s announcement centers on whether a large open-weight model can provide useful coding and agentic performance at a lower inference cost. Reflection says Beam scores comparably to GLM 5.2 on advanced reasoning benchmarks while using three to four times less inference compute. It also says models in the 2-trillion-parameter-plus range, such as Qwen 3.8-Max, require significantly more compute per token.
That could matter to organizations running coding assistants or agents at scale, where compute needs affect operating costs and deployment choices. The comparison is based on estimates, however, rather than a complete measure of real-world serving costs. Reflection’s analysis estimates generation compute from active parameters and generated tokens and excludes prompt prefill, context-dependent attention operations and serving overhead.
Open weights would let developers inspect and run the model under the terms that Reflection publishes. Until the weights and license are available, users cannot assess those practical options directly. The model’s performance and efficiency claims also remain company-reported results awaiting scrutiny through the forthcoming technical materials and further evaluations.
As an affiliate, we earn on qualifying purchases.
Training at Reinforcement Learning Scale
Beam is Reflection’s first open-weight model. The company says it trained the model with a focus on coding and agentic performance, using a sparse Mixture-of-Experts design in which 23 billion of 501 billion total parameters are active per token. That distinction is relevant to Reflection’s compute comparisons: its estimates use active parameters rather than total model size.
Reflection describes large-scale reinforcement learning as a central part of Beam’s development. Its account says the training effort generated more than 100 million rollouts, with a maximum context length of 256,000 tokens. The company also says it developed methods to manage asynchronous training, including cases where training samples came from model versions more than a day old. These are descriptions of the company’s approach; the public announcement does not yet include the full technical report needed to examine them in detail.
Reflection’s benchmark chart compares Beam with a selection of open models across coding, terminal use, reasoning and STEM tasks. It identifies some scores as not reported, and the company acknowledges that models such as Kimi K3 remain ahead on raw capability. Reflection’s stated case for Beam is instead its balance of coding and agentic results with inference efficiency.
““Beam’s advantage is efficiency at inference time.””
— Reflection
large language model training datasets
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Release and Evaluation Details Pending
Beam’s weights are not yet public, so developers cannot independently reproduce the reported results or inspect the model. Reflection has not yet published the technical report, model card, developer artifacts or licensing terms in the supplied announcement. Those materials will clarify the model’s limits, recommended use and conditions for access.
The company says final red-teaming and evaluations are underway, but has not described their scope or published findings. It is also unclear whether the benchmark results will be reproduced by outside evaluators, how performance will vary across real-world coding tasks, or how much deployed inference will cost after serving overhead and prompt processing are included.
GPU high-performance computing for AI
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Weights Planned Later This Month
Reflection says it plans to release Beam’s weights, technical report, model card and developer artifacts later this month. The company is accepting sign-ups for early access while final red-teaming and evaluations continue. No specific release date is provided in the announcement.
Once those materials are available, developers and independent evaluators will be able to examine the benchmark methodology, test the model and review the access terms. Until then, the reported scores and compute comparisons remain Reflection’s claims, and the model’s public availability remains pending.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is Beam?
Beam is Reflection’s first open-weight model, a sparse Mixture-of-Experts system intended for coding, reasoning and agentic workloads. Reflection reports 501 billion total parameters, with 23 billion active per token.
Can developers download Beam now?
No. Reflection says the model is undergoing final red-teaming and evaluations. It plans to release the weights and related technical materials later this month, but has not given a specific date.
What performance has Reflection reported?
Reflection reports scores including 80.1 on Terminal Bench 2.1 and 80.9 on SWE-bench Verified. It says Beam is competitive with larger open models on some coding and agentic tasks, while acknowledging that some models remain ahead on raw capability.
What does Reflection say about Beam’s compute efficiency?
The company says Beam achieved reasoning scores comparable to GLM 5.2 using three to four times less inference compute. Its estimate excludes prompt prefill, context-dependent attention operations and serving overhead, so it is not a complete measure of deployment cost.
Source: hn
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
