Short answer: Beam is Reflection AI's first open-weight model, announced on October 5, 2026. It is a sparse mixture-of-experts model with 501 billion total parameters and 23 billion active per token, built for coding, reasoning and agentic work. Today it is available only through an early-access waitlist; Reflection says the weights, technical report and model card will ship later in October under an Apache 2.0 license (Reflection).

Reflection has been one of the most closely watched open-model labs in the US, mostly because it had not shipped anything. Beam is its first release, and the company frames it as a push for the "Western open-weight frontier" against strong Chinese open models such as GLM, Kimi, Qwen and DeepSeek. This guide covers what Beam is, how it performs on the published numbers, where critics push back, and what it means if you build software with AI. For the wider map of assistants, agents and models, start with our AI tools hub.
What is Reflection Beam?
Beam is a text-only large language model from Reflection AI, a startup founded in 2024 by former Google DeepMind researchers Misha Laskin and Ioannis Antonoglou and backed by Nvidia, according to Reuters.
The key design choice is sparsity. Beam has 501 billion parameters in total, but only 23 billion work on each token it generates. A mixture-of-experts model routes each token to a small set of specialist sub-networks instead of running the whole network every time. That keeps the knowledge of a very large model while cutting the compute needed per answer.
Reflection says it trained Beam with a particular focus on coding and agentic tasks: multi-step work where the model plans, calls tools, runs commands and checks results. The company calls it "the first model in a series" and says it is already training the next one.
How was Beam trained?
Beam was trained in two big phases, and Reflection published unusually detailed numbers for both.
- Pretraining: 23.8 trillion curated tokens from the web, public sources and licensed datasets, on a cluster of 6,144 NVIDIA GB300 NVL72 GPUs, finished end to end in under four weeks.
- Reinforcement learning (RL): more than 100 million rollouts on 10,500 NVIDIA GB300 GPUs over four weeks, with a maximum context of 256,000 tokens per rollout and roughly 1.3 billion sandboxes used for training and grading.
The RL phase drew on a pool of nearly one million environments covering software engineering, terminal use, competitive coding, STEM, web search and tool use. Reflection says capabilities kept improving as RL compute grew, "with no sign of a plateau."
One detail matters for developers. Reflection trained Beam with a length penalty that rewards correct answers while discouraging wasted tokens. The result is exposed as a reasoning effort setting, so you can trade speed for depth on each request.
How does Beam perform on coding benchmarks?
Beam is competitive with mid-size open models but behind the strongest Chinese ones on raw coding scores. These are Reflection's own reported results, not independent tests:
| Benchmark | Beam | GLM 5.2 | Qwen 3.8 Max | Kimi K3 | DeepSeek V4.1 Flash |
|---|---|---|---|---|---|
| DeepSWE v1.1 | 44.4 | 44.0 | 51.0 | 68.0 | 74.2 |
| Terminal Bench v2.1 | 80.1 | 81.0 | 86.6 | 88.3 | 90.6 |
| SWE Bench Pro v1 | 65.5 | 62.1 | 67.7 | not reported | not reported |
On SWE-bench Verified, Beam reports 80.9, and on SWE-bench Multilingual it reports 78.0. Reflection itself says Kimi K3 "remains ahead on raw capability."
So the pitch is not "best coding model." The pitch is efficiency. Reflection claims Beam reaches reasoning scores comparable to GLM-5.2 while using three to four times less inference compute. For context, Reuters notes that GLM-5.2 has 744 billion total parameters and 40 billion active, against Beam's 501 billion and 23 billion.
Is the efficiency claim solid?
It is plausible but approximate, and Reflection says so in its own fine print. The company estimates compute as roughly two times active parameters times generated tokens. It states that these estimates exclude prompt prefill, attention costs that grow with context and serving overhead, so they are "an approximate compute comparison rather than measured inference cost."
Critics have focused on exactly that gap. Turing Post called the release disappointing, pointing to the 30-point gap with DeepSeek V4.1 Flash on DeepSWE and arguing that the efficiency numbers do not let you compare real cost per task.
The fair reading: Beam's per-token compute is low because only 23 billion parameters are active. Whether that turns into a lower bill depends on the price providers charge, how long its reasoning runs on your tasks and how much context you send. Nobody can measure that until the weights and API pricing are public.
How can you use Beam today?
Right now, only through early access. Reflection is giving the preview to a select group of users via a waitlist while the model goes through final red-teaming and evaluations.
The developer documentation already describes how it will work (Reflection docs):
- Model ID:
Beam-501B-A23B. - API: OpenAI-compatible, so existing client code can point to Reflection's endpoint with a new base URL and key.
- Context: 256,000 tokens in the API (marked beta), with up to 128,000 output tokens. The blog says midtraining extended the model's effective context to 1 million tokens, so the API limit may change.
- Reasoning effort:
low,medium,high,xhighormax, withmediumas the default. Reasoning cannot be switched off. - Features: tool calling and structured outputs.
- Knowledge cutoff: June 30, 2026.
Reflection has not published API pricing yet. Once the weights land under Apache 2.0, you will be able to run, fine-tune and ship Beam in commercial products, as long as you keep the license and notices. Reflection says it will launch with distribution partners and integrations with open-source libraries and agent harnesses.
Can you run Beam on your own hardware?
Not on a laptop. Even with only 23 billion active parameters, all 501 billion have to sit in memory. At 16-bit precision, the weights alone take roughly 1 TB before any quantization. That puts Beam in multi-GPU server territory.
This is where open weights pay off for companies, not hobbyists. Reflection says it can serve its models through its API or let you deploy them in your own environment, including private cloud, on-premises, air-gapped and edge setups. For teams in regulated sectors that cannot send code to a third-party API, a strong open model they can host themselves is the real value.
What can Beam actually build?
Reflection published four demos that show the agentic side of the model:
- A live NYC subway map: Beam found the public MTA data, checked authentication, built the frontend and backend, and kept the server running for live updates.
- A 3D browser game in p5.js: an astronaut falling toward Earth who dodges or destroys asteroids. Beam is text-only but reasoned about the visuals in code.
- A fine-tuning notebook: plugged into the OpenCode agent, Beam read the Unsloth docs and built a notebook to fine-tune the smallest Gemma-4 model on a Text-to-SQL task. Reflection reports a 66.5% accuracy gain on the held-out test set.
- A geography generalization test: a 16,200-point land-or-water grid, where Beam scored 95.5% coverage.
Reflection also reports that during RL on coding and terminal tasks, Beam improved at browsing without being trained on it, and learned to query other models and use OCR APIs when given web access. Treat demos as demos: they show what is possible, not what you will get on every run.
How does Beam fit in the coding agent landscape?
Beam is a model, not a coding tool. You do not "open Beam" the way you open Cursor or Claude Code. You plug it into an agent harness, an IDE extension or your own pipeline, the same way teams use GLM, Qwen or DeepSeek today.
That makes the choice of agent as important as the model. Our comparison of the best AI coding agents explains which harnesses accept custom or self-hosted models, and the AI tools for developers map shows where models, agents and editors connect.
For most developers, the practical question is simple: when the weights ship, will Beam beat the open model you already use, at a cost you can afford, on your own tasks? Run your own evaluation on a slice of your real work before switching.
The skill behind Beam
Every month brings a new open model with a new benchmark table. The tools change fast; the skill that lasts is knowing how to evaluate them. That means reading a benchmark table critically, building a small eval from your own tasks, wiring a model into an agent through an OpenAI-compatible API and reviewing the code it writes before it reaches production.
That is the work of an AI engineer: building with models like Beam and checking what they produce, not just prompting them. If you want to learn that path, compare our programs and see which one fits where you are today.
