October 6, 2026Open SourceRLCoding

Reflection's Beam: 501B Parameters, and an Honest Scoreboard

Reflection AI has raised about $4.7 billion and until Monday had not shipped a frontier model. Now there is one. Beam is a 501-billion-parameter mixture-of-experts model with 23 billion parameters active per token, text-only, with a 1M-token context, aimed at coding and agentic work. Weights, a technical report and a model card are promised for later this month. Today there is an early-access signup and a long blog post.

The interesting part is how it was trained. Pretraining took 23.8 trillion tokens and under four weeks on 6,144 Nvidia GB300 GPUs. Then came the reinforcement learning run: more than 100 million rollouts over four weeks on 10,500 GB300s, about 1.3 billion sandboxes for training and grading, and a million coding, agentic and STEM environments. Reflection says scores kept climbing with RL compute and showed no plateau. For scale, the post says Thinking Machines' Inkling used 30 million rollouts.

The scoreboard is unusually candid for a launch post. On Terminal-Bench 2.1 Beam scores 80.1, next to GLM 5.2 at 81.0, and behind Kimi K3 at 88.3 and DeepSeek V4.1 Flash at 90.6. On DeepSWE it gets 44.4 against 68.0 for K3 and 74.2 for DeepSeek. It does beat Inkling wherever both report numbers, for example 65.5 against 54.3 on SWE-Bench Pro v1. Reflection says outright that the top Chinese open models remain ahead on raw capability.

So the pitch is efficiency. With 23B active parameters against roughly 40B for GLM 5.2, Reflection claims comparable reasoning scores at 3 to 4 times less inference compute. Those numbers are self-reported FLOP estimates and nobody outside has checked them. TechCrunch notes the company has locked up more than $7 billion of GB300 capacity through 2029 from SpaceX and Nebius.

Two takeaways. The best Western open-weight model now openly benchmarks itself as a chaser of Chinese labs, which would have read as satire two years ago. And the recipe is turning into a formula: a mid-size sparse model plus an enormous RL run across a million agent environments. The environments, more than the parameter count, are where the money went.

Link: reflection.ai/blog/introducing-beam
← Previous
Wikimedia Found OpenAI's Rogue Agents in Its Own Backyard
Next β†’
A Second Lab's Agents Showed Up on urlquery, and This Fleet Looks Like Tencent's
← Back to all articles

Comments

Loading...
>_