
Tech • AI • Robotics • Game
Reflection’s new open model Beam gives the United States its first broadly competitive answer to China’s leading downloadable AI systems, even as Chinese labs still lead on top benchmarks and Mistral pushes Europe back into the race.
Open models handled 56% of all AI traffic on Vercel’s AI Gateway in August, up from under 10% in December. Over the past year, Chinese models accounted for about 41% of downloads on Hugging Face, reflecting how labs such as DeepSeek, Qwen, Kimi, and GLM came to dominate freely available high-end AI.
New York startup Reflection, founded in 2024 by former Google DeepMind researchers Misha Laskin and Yanis Antaglu, released Beam, a 501 billion-parameter mixture-of-experts model. Although only 23 billion parameters are active per token, the company trained it on 23.8 trillion tokens of text and code, then ran a four-week reinforcement-learning phase on 10,500 Nvidia GB300 chips with more than 100 million task attempts inside 1.3 billion isolated sandboxes.
On DeepSWE, a software-bug benchmark, Beam scored 44, matching GLM 5.2 but behind GLM 5.3 at 61, Kimi K3 at 68, and DeepSeek V4.1 Flash at 74. On Humanity’s Last Exam, Beam scored 36 against 47 for Kimi K3. The result places Beam roughly around China’s previous generation rather than the current frontier.
Reflection argues Beam’s value is not raw top score but cost-performance. The company says Beam can match GLM 5.2 on hard reasoning while using three to four times less compute, and widened that advantage against very large models such as Qwen 3.8 Max. Beam was explicitly rewarded for concise, useful answers rather than long responses, and it offers adjustable inference effort for speed or depth.
Reflection said about 110,000 attempts were running simultaneously during Beam’s reinforcement-learning phase, while the model kept updating without destabilizing training. Engineers reported 71 failures during the run without halting it, added separate judging systems to reduce reward hacking, extended context to 1 million tokens, and observed Beam improve at web use and tool calling even without direct browsing-specific training.
Beam is due to be released under Apache 2.0, allowing commercial use, with a stronger successor already in training. Reflection, backed by Nvidia, Sequoia, and CRV, was valued at $25 billion in June and has secured compute through deals including SpaceX and Nebius. Its pitch is aimed at companies and governments that want open models but are unwilling or unable to rely on Chinese systems.
A day after Beam’s unveiling, Mistral introduced Magistral Large 4, a 1.05 trillion-parameter model with 52 billion active parameters, image capabilities, 1 million-token context, and support for more than 160 languages. It was trained on roughly 3,800 Nvidia chips in European data centers and is priced at $1.36 per million input tokens and $4.18 per million output tokens ahead of a wider release.
Mistral said its new model outperformed DeepSeek V4 Pro and Qwen 3.8 Max on its combined coding score and ranked second only to Claude Opus 5 in a blind professional coding review. The company also highlighted security performance: 82% on a vulnerability recreation-and-fix test, 93% solved on 40 hacking-style challenges, and 93.3% blocked prompt-injection attacks, while still refusing overtly malicious requests more often than other open models.
Reflection argues the United States lost momentum in open models after Meta’s Llama 3, while Chinese labs treated open release as a strategic priority and diplomatic tool. Mistral is making a similar sovereignty argument for Europe, saying AI can be trained, run, and deployed entirely under European law. The result is an increasingly regional contest over who supplies the most capable open systems to enterprises and governments.
Beam does not yet displace the strongest Chinese open models, but it gives the United States a credible, efficient entrant in a fast-growing market. With Mistral also scaling up, the next phase of open AI competition is shifting from novelty to infrastructure, cost, and political trust.
Ask a question