Reflection Beam Explained: Specs, Benchmarks & When You Can Download It
Status at a glance (October 6, 2026): Beam has been announced but the weights are not yet released. Every benchmark score below comes from Reflection itself and has not been independently verified. We will update this page when either of those changes.
What is Reflection Beam?
In one sentence: Beam is a large, open-weight AI model from the US startup Reflection AI, designed for writing code, reasoning through hard problems and acting as an “agent” that carries out multi-step tasks.
“Open-weight” means that once it is released, anyone can download the model’s trained parameters and run it on their own hardware — unlike ChatGPT or Claude, which you can only use through the company’s own service.
Reflection AI was founded in 2024 by Misha Laskin and Ioannis Antonoglou, both former researchers at Google DeepMind (Antonoglou worked on AlphaGo and AlphaZero). The company is backed by Nvidia and, according to press reports, has raised money at a valuation of around $25 billion. Beam, announced on October 5, 2026, is its first public model.
The bigger story: over the past two years, the best open-weight models have mostly come from Chinese labs — DeepSeek, Alibaba’s Qwen, Moonshot’s Kimi and Z.ai’s GLM. Beam is pitched as a Western answer to that wave.
Key specs
- Size: 501 billion total parameters, about 23 billion active per token.
- Architecture: sparse mixture-of-experts (MoE). Think of it as a large team of specialists where only a few are called on for each word — you get the knowledge of a big model at a fraction of the running cost.
- Input/output: text only (no images or audio).
- Context window: up to 1 million tokens, according to Reflection (reinforcement-learning training used up to 256K).
- Training data: 23.8 trillion tokens of pretraining, trained from scratch (not a fine-tune of someone else’s model).
- Focus: coding, reasoning and agentic (tool-using) tasks.
- Efficiency claim: Reflection says Beam reaches reasoning scores comparable to GLM-5.2 while using 3–4× less inference compute. The company itself calls this an approximate comparison, not a measured cost.
- License: Apache 2.0 (permissive, commercial use allowed) — once the weights ship.
Beam vs. the Chinese open models
Reflection’s own summary is refreshingly modest: Beam is “competitive with” GLM-5.2, “approaching” Qwen 3.8-Max, and behind Moonshot’s Kimi K3 on raw capability. Its selling point is doing a lot with less compute.
Here is a selection of rows from Reflection’s announcement. Higher is better; “—” means the score was not reported.
| Benchmark | Beam | GLM-5.2 | Qwen 3.8-Max | Kimi K3 | DeepSeek V4.1 Flash |
|---|---|---|---|---|---|
| Terminal Bench v2.1 (agentic coding) | 80.1 | 81.0 | 86.6 | 88.3 | 90.6 |
| DeepSWE v1.1 (coding) | 44.4 | 44.0 | 51.0 | 68.0 | 74.2 |
| SWE Bench Pro v1 (coding) | 65.5 | 62.1 | 67.7 | — | — |
| Humanity’s Last Exam, no tools (reasoning) | 36.2 | 40.5 | 43.6 | 46.9 | 39.1 |
| GPQA Diamond (science Q&A) | 90.5 | 91.2 | 92.6 | 93.5 | 90.9 |
| MCP Atlas (tool calling) | 78.7 | 77.8 | 84.5 | 82.3 | — |
| IFBench (following instructions) | 79.7 | 73.3 | 82.8 | — | — |
Source: Reflection’s Beam announcement, October 5, 2026. Vendor-reported, not independently verified.
How to read this:
- Against GLM-5.2 it is roughly a tie — ahead on some coding and tool-use rows, behind on hard reasoning (Humanity’s Last Exam). GLM-5.2 is reported to be a larger model (around 750B total parameters), which is why the efficiency angle matters.
- Against Qwen 3.8-Max, Beam trails on almost every row — but Reflection notes Qwen 3.8-Max has over 2 trillion parameters, roughly four times Beam’s size.
- Against Kimi K3 and DeepSeek V4.1 Flash, Beam is behind on most reported rows, sometimes by a wide margin (e.g. DeepSWE).
So the honest summary is: Beam does not beat the best Chinese open models. Its claim is that it gets reasonably close while being cheaper to run — and that it is a strong, permissively licensed option built in the US.
How to try or download Beam
Right now (October 6, 2026), you can’t download it. Reflection says the model is in final red-teaming and evaluation. What exists today:
- Early access waitlist on Reflection’s platform (
platform.reflection.ai). - Weights “later this month” (October 2026) under Apache 2.0, along with a technical report, a model card and tooling for running, evaluating and fine-tuning the model.
A practical note: even with only 23B active parameters, all 501B parameters still have to fit in memory. Running Beam yourself will realistically mean multi-GPU server hardware, not a laptop — unless smaller quantized versions appear from the community after release.
We will update this section with download links and hardware requirements once the weights are out.
FAQ
What is Reflection Beam?
Who made Beam?
How big is Reflection Beam?
How does Beam compare to DeepSeek?
When can I download Beam?
Can I trust Beam’s official benchmarks?
Is Reflection Beam open source?
Who is Beam for?
Bottom line
Beam is a credible, efficiency-focused open model from a well-funded US lab — not a Kimi or DeepSeek killer, and Reflection doesn’t claim it is. The two things to watch this month: the actual weight release, and independent benchmark results. Until both arrive, treat every number on this page as the company’s own claim.
Want to turn benchmark numbers like these into a clean chart for a slide or post? Paste them into our free bar chart maker — no sign-up, nothing uploaded.