Reflection AI Beam Locally or on AWS: Specs, Benchmarks, Pricing, and How to Run It

Logeshwaran
—

Reflection AI Beam is the first open-weight model from Reflection AI, the Nvidia-backed US lab. It was announced on October 5, 2026: a 501-billion-parameter mixture-of-experts model that uses 23 billion parameters per token, built for coding and agent work, and promised under the Apache 2.0 license. Two things surprise people once they look past the headlines. First, you cannot download it yet. The weights arrive "later this month"; today Beam is a waitlisted beta API. Second, the headlines call it America's answer to Chinese open models, but Reflection's own launch chart puts Beam behind Kimi K3, GLM-5.3 and DeepSeek V4.1 Flash on most of the coding tests where they share a score. What Beam actually sells is efficiency: roughly GLM-5.2-level answers for about a third to a quarter of the compute. Below: every spec, what the benchmark chart really says, how to get API access today, what "free" means here, and the honest hardware math for running Beam locally or on AWS once the weights land.

Jake saw the news on his phone during a slow afternoon at the shop: "US lab unveils open model to take on China." He had been paying for an AI subscription to write repair quotes and product listings, and "open" sounded like "free and mine." By closing time he had a plan. Download Beam, put it on the shop PC, cancel the subscription. He texted Ethan a screenshot of the headline with three fire emojis. Ethan replied with one question: "How much memory does the shop PC have?" Jake checked. Thirty-two gigabytes. "Then let's talk before you cancel anything," Ethan wrote. "Beam is real, and it might be very good. But even a heavily compressed copy needs about eight times the memory in that PC, and right now nobody can download it anyway."

⚡ Quick Answer

• What it is → a 501B-total, 23B-active open-weight model for coding and agents, text only. Full specs.

• Can you download it? → not yet. Weights are promised later in October 2026 under Apache 2.0. Use the API today.

• Is it free? → the weights will be; the API has no public price yet. Pricing explained.

• Run it locally? → not on a normal PC. A 4-bit copy needs about 250 to 280 GiB of memory. The hardware math.

Running it in the cloud? The AWS instances that fit, with real hourly prices.

🧭 NEW HERE? READ THESE FIRST

New to open models, local AI or AWS? These five pages cover the basics this one builds on:

📌 Bookmark this; Beam's weights are due later this month, and this page will have the release-day details.

What Reflection AI Beam is, in plain English

A few words do most of the work in every Beam headline, and each one is easy to misread. Here they are in plain English before we go any deeper.

  • A model is the trained AI itself: a very large file of numbers, called weights, plus the code that runs them. Chat apps and coding tools are just windows onto a model.
  • Open-weight means the company publishes those weights so anyone can download them, run them on their own hardware, and adapt them. It is not quite the same as open source: the training data and most of the training code usually stay private. Beam will be open-weight. Reflection says the release includes the weights, a technical report, a model card, and "the full stack for running, evaluating, and fine-tuning the model."
  • Apache 2.0 is one of the most permissive licenses in software. In practice it means you can use the model commercially, modify it and ship products on it, as long as you keep the license and notices. That is more generous than the custom licenses on Kimi K3 and GLM-5.3.
  • Parameters are the individual numbers in the model. More parameters usually means more knowledge and skill, and always means more memory.

What "501B-A23B" means: a mixture of experts

Beam's model ID on Reflection's API is Beam-501B-A23B. The first number is the total size: 501 billion parameters. The second, after the "A," is how many are active for each word the model produces: 23 billion. That is about 4.6% of the model working at any one moment.

That gap exists because Beam is a mixture-of-experts (MoE) model. Inside, instead of one giant block that processes every word, there are many smaller "expert" blocks and a router that picks a few of them for each token. Reflection describes Beam as using "fine-grained routed experts," which means many small experts rather than a handful of big ones.

"Think of a hospital," Ethan told Jake. "It has hundreds of doctors on the payroll, but when you walk in with a broken wrist, you only see three of them. The hospital still needs a building big enough for all of them, though."

That analogy carries the single most important fact for anyone hoping to run Beam: the 23 billion active parameters decide how fast it runs, but the 501 billion total decide how much memory it needs. All the experts have to sit in memory, because the router might call any of them for the next word. A 23B-active model does not fit where a 23B model fits. We come back to this in the hardware section.

Who is Reflection AI?

Reflection AI was founded in 2024 by two former Google DeepMind researchers, Misha Laskin and Ioannis Antonoglou, and is based in Brooklyn, New York. Its first product was Asimov, a code-comprehension agent that reads a company's source code, documents and messages to answer questions about how its software is built. Beam is its first foundation model.

The company has raised money at a remarkable pace. It came out of stealth in March 2025 with $130 million, raised $2 billion at an $8 billion valuation in October 2025 in a round led by Nvidia, and reached a valuation of about $25 billion in 2026. It has also signed large compute deals, including access to Nvidia GB300 systems at SpaceX's Colossus data center and a separate agreement with Nebius, plus a partnership to build a 250-megawatt data center in South Korea with Shinsegae, and a role as a model provider to the US Department of Energy's national laboratories. Its stated mission is to build frontier open intelligence in the United States, and Beam is the first public proof of that plan.

If you came here because you searched "does Reflection AI have a model," the answer changed on October 5: yes, and this is it. And if you are thinking of a different Reflection, the 2024 "Reflection 70B" fine-tune of Llama was an unrelated project by a different team. It has nothing to do with Reflection AI or Beam.

Reflection AI Beam specs at a glance

Everything Reflection has published about Beam so far, in one table. Two rows deserve a second look: the context window, which depends on where you read it, and the "text only" line, which matters if you hoped to send it screenshots.

SpecBeam
MakerReflection AI (Brooklyn, New York)
AnnouncedOctober 5, 2026
Model ID (API)Beam-501B-A23B (capitalization matters)
ArchitectureSparse mixture of experts, 52 layers, interleaved local and global attention
Parameters501 billion total, 23 billion active per token
Context window1 million tokens effective after training; 256K on the beta API (input and output combined), "may change during the beta"
Max output128K tokens per response on the API
Knowledge cutoffJune 30, 2026
Input and outputText only; no images, audio or files
ReasoningAlways on; effort low, medium (default), high, xhigh, max
Also supportsTool calling, structured JSON outputs, streaming
Pretraining23.8 trillion tokens, under four weeks on 6,144 Nvidia GB300 GPUs
Reinforcement learning10.5K GB300 GPUs for four weeks, over 100 million rollouts, up to 256K-token rollouts
LicenseApache 2.0 (promised with the weight release)
WeightsNot released yet; due "later this month" (October 2026)
Access todayWaitlisted beta API at platform.reflection.ai

The training story in brief

Reflection was unusually open about how Beam was made, and a few details help explain what kind of model it is. About 95% of the raw internet text it collected was thrown away during cleaning, and the company says its own quality filters kept roughly 1.8 trillion good tokens that common filters would have discarded, including most of its curated code from the web. Code and technical writing were deliberately repeated during training to deepen the model's grounding in programming. The pretraining run itself was smooth by industry standards: nine semi-automatic rewinds, and 92.3% of wall-clock time spent on training steps that made it into the final model by the end.

The reinforcement learning phase, where the model practices tasks and is rewarded for solving them, was huge: nearly one million training environments covering software engineering, terminal use, competitive coding, science, web search and tool use, with up to 170,000 sandboxes running at once and about 1.3 billion sandboxes used in total. Reflection says capability was still climbing when the run ended. A separate safety-focused model was trained from the same base, and the two were merged through a distillation step. The safety evaluation results have not been published yet; Reflection says they will come with the technical report.

One training detail shapes how Beam feels to use: Reflection trained it with a length penalty that rewards correct answers reached with fewer tokens. That is where the reasoning-effort setting comes from. At low, Beam thinks briefly and answers fast. At max, it is allowed to think for a long time on hard problems. Because reasoning is mandatory, even low spends some tokens thinking.

Beam benchmarks: what Reflection's own chart really shows

Reflection published a detailed chart comparing Beam with seven other models. It deserves credit for including models that beat Beam, and for marking missing scores as "not reported" instead of hiding them. Here are the rows that matter most, copied from that chart. Higher is better on every row, and "NR" means not reported.

BenchmarkBeamGLM 5.2GLM 5.3Kimi K3Qwen 3.8 MaxDeepSeek V4.1 Flash
Terminal Bench v2.180.181.088.288.386.690.6
SWE Bench Pro v2-Hard77.2NR84.388.2NRNR
SWE Bench Pro v165.562.1NRNR67.7NR
DeepSWE v1.144.444.061.068.051.074.2
SWE Atlas Codebase QnA34.6NR61.068.0NRNR
AIME 2026 (math)97.899.2NRNRNRNR
HLE, no tools36.240.542.346.943.639.1
GPQA Diamond (science)90.591.291.793.592.690.9
MCP Atlas (tool use)78.777.884.282.384.5NR
AutomationBench public37.026.248.246.739.854.8
IFBench (following instructions)79.773.3NRNR82.8NR
AA-LCR (long context)79.378.379.788.780.384.0

Read across the rows and a clear picture forms.

  • Against GLM-5.2, Beam is roughly even. It wins on SWE Bench Pro v1, DeepSWE (barely), MCP Atlas, AutomationBench, instruction following and long context. It loses on Terminal Bench (by less than a point), math, HLE and GPQA. That matches Reflection's own words: "competitive with larger open models like GLM 5.2."
  • Against Kimi K3, GLM-5.3 and DeepSeek V4.1 Flash, Beam trails on almost every shared row. The gaps are not small on some coding tests: 34.6 against 61.0 and 68.0 on codebase questions, 44.4 against 74.2 on DeepSWE. Reflection says so plainly: "frontier open models like Kimi K3 remain ahead on raw capability."
  • Against the other Western open models, Beam leads. The chart also includes Inkling and Nvidia's Nemotron 3 Ultra, and Beam beats both on nearly every row they share, for example 80.9 against 77.6 and 70.7 on SWE Bench Verified, and 80.1 against 63.8 and 56.4 on Terminal Bench. That is the honest version of "America's best chance": the strongest Western open-weight model in this chart, not the strongest open model overall.

The DeepSeek comparison is the most telling. DeepSeek V4.1 Flash activates only 8 to 16 billion parameters per token by its own model card, fewer than Beam's 23 billion, and its API costs $0.15 per million input tokens off-peak. In this chart it scores higher than Beam on every coding row where both have a number. Beam's efficiency story is real against GLM-5.2, but it is not unique.

The "3 to 4 times less compute" claim, explained

Reflection's headline efficiency number says Beam reaches GLM-5.2-level scores on advanced reasoning tests "while using 3–4× less inference compute." The method behind it is simple, and worth understanding before you repeat it. Reflection estimated compute as 2 × active parameters × the average number of tokens generated per attempt. Two things push that number down for Beam: fewer active parameters than GLM-5.2, and shorter answers, thanks to the length penalty in training.

Reflection is also upfront about what the estimate leaves out: reading the prompt (prefill), attention costs that grow with context length, and real serving overhead. So it is a fair comparison of how much "thinking" each model does per answer, not a measured bill. A real cost comparison needs real prices, and Beam does not have a public price yet.

"So it's a fuel-economy sticker," Jake said, "not a receipt."

"Exactly," Ethan said. "A useful sticker. But you only know your real cost when you've driven it on your own roads."

Every number here is Reflection's own, for now

Independent leaderboards such as Artificial Analysis had no Beam entry on launch day, and outside researchers cannot test a model they cannot download. Benchmarks from a model's maker are not worthless, and Reflection's chart is more honest than most because it shows its losses, but they are always run under the maker's best settings. Treat the chart as a starting point, and wait for independent runs before you bet a product on Beam.

Reflection did share a few demos alongside the chart. Beam built a live New York City subway map from public transit data, a 3D browser game, and a fine-tuning notebook for Google's smallest Gemma 4 model that improved its held-out accuracy by 66.5%. On a viral "is this point land or water" grid puzzle, Beam got 95.5% coverage, which Reflection places between Opus 5 (92.5%) and Fable 5 (97.8%). Demos show what is possible on a good day; the chart shows the average day.

How to use Reflection AI Beam today: the waitlist and the API

Until the weights are released, the only way to use Beam is Reflection's own API, and that is in a waitlisted beta. Here is the whole path, from sign-up to a first answer.

  1. Join the waitlist. Sign up at platform.reflection.ai. Reflection is giving access to "a select group of users" first, so there may be a wait. API keys become available once your account is enabled.
  2. Create a project and an API key. In the platform, open API Keys, pick a project, and create a key. Copy it straight away; the full key is shown only once. Keys belong to both a project and the person who made them, and a key stops working if its creator leaves the organization.
  3. Add a payment method if asked. Some organizations must verify a credit card in Billing before keys work. If your first request returns 403 payment_method_required, that is the step you skipped.
  4. Save the key as an environment variable called REFLECTION_API_KEY, never in your code or a public repository.
  5. Send your first request to the OpenAI-compatible endpoint, shown below.

A first request with curl, which works in Windows Terminal, macOS Terminal and Kali alike:

curl https://api.reflection.ai/openai/v1/chat/completions \
  -H "Authorization: Bearer $REFLECTION_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Beam-501B-A23B",
    "reasoning_effort": "low",
    "max_completion_tokens": 4000,
    "messages": [
      {"role": "user", "content": "Write a two-sentence repair quote for a cracked phone screen."}
    ]
  }'

And the same thing in Python, using the official OpenAI library pointed at Reflection:

import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.reflection.ai/openai/v1",
    api_key=os.environ["REFLECTION_API_KEY"],
)

reply = client.chat.completions.create(
    model="Beam-501B-A23B",
    reasoning_effort="medium",
    max_completion_tokens=8000,
    messages=[{"role": "user", "content": "Explain RAID 1 to a shop owner in three sentences."}],
)
print(reply.choices[0].message.content)

Jake's first test was exactly that kind of everyday job, a repair quote, and Beam handled it in seconds at low effort. The model is built for much harder work than that, but small tasks are the right way to learn its habits before you trust it with anything bigger.

Reasoning effort: the setting that controls cost and speed

Beam always reasons before it answers, and you choose how much with reasoning_effort. The accepted values are low, medium, high, xhigh and max. If you leave it out, Beam uses medium. The OpenAI values none and minimal are rejected with a 400 unsupported_value error, because there is no way to switch Beam's thinking off.

Two practical consequences follow. First, reasoning tokens are billed and rate-limited like any other output, and they count toward max_completion_tokens. Second, if you set that limit too low, Beam can spend the whole budget thinking and return finish_reason: "length" with an empty answer. When that happens, raise the limit or lower the effort. Start at medium, drop to low for simple, latency-sensitive jobs, and go up only when answers are not good enough.

The model's thinking comes back separately, in a field called reasoning_content, while the answer stays in content. Show users the answer, not the reasoning. When you build tool-calling loops, send the assistant's message back unchanged, including its reasoning_content, so the model keeps its train of thought.

What works, what doesn't: the OpenAI-compatibility gotchas

Reflection's endpoint copies OpenAI's Chat Completions format, so most OpenAI-based code works after changing the base URL and key. "Most" is doing some work, though. These are the differences most likely to bite:

  • Only Chat Completions and Models exist. The Responses, Embeddings, Images, Audio, Files, Batch and Assistants endpoints are not supported. A tool built on the newer Responses API will not work with Beam.
  • stop sequences are accepted but ignored. Generation does not stop at them. If your code relies on stop strings to cut output, trim the text yourself. This one fails silently, so it is worth knowing before you debug for an hour.
  • Text only. Image, audio and file inputs are rejected.
  • Some parameters are fixed: n must be 1, logprobs must be false, store must be false (completions are not stored), and logit_bias must be empty.
  • Unknown parameters fail loudly. Fields such as user, metadata, top_logprobs or prediction return 400 unsupported_parameter. If a framework sends one by default, turn it off in that framework's settings.

The good news is the list of things that do work: system and developer messages, streaming with usage reports, tool calling with tool_choice and parallel calls, and structured JSON outputs with or without a schema.

Use Beam inside a coding agent: Mirror CLI, OpenCode, Pi and Hermes

Beam was trained for agentic coding, so most early users will meet it inside a coding agent rather than a chat box. Reflection documents four routes:

  • Mirror CLI is Reflection's own terminal coding agent, in early beta. It uses Beam by default, so setup is mostly pasting your API key. If you have used Mirror with another provider, start it with mirror --api reflection --model Beam-501B-A23B, and set effort with --reasoning-effort high or the /reasoning command.
  • OpenCode works through a custom provider entry in its config file. Give the model a context limit of 262,144 and an output limit of 131,072; without the limit, OpenCode never compacts a long session. Its effort picker offers low, medium and high, and you add xhigh and max as variants.
  • Pi works the same way, with its thinking levels mapped to Beam's five effort levels.
  • Hermes works too, with one known quirk: its automatic session-title request asks for no reasoning, Beam refuses, and a one-shot run can exit with "Auxiliary title generation failed: Connection error." It is harmless.

Two habits save trouble in every agent. Check that the model ID is exactly Beam-501B-A23B, capital letters included, and remember that an agent fires many requests per task, all counted against your organization's rate limits.

Rate limits and the daily allowance

Beta limits apply to your whole organization, shared across every project and key, so creating more keys does not raise them. Limits are counted in requests and tokens, per minute and per UTC day, and there is also a cap on how many requests can be in progress at once. Your plan may add a daily token allowance. Daily limits reset at 00:00 UTC, which is 5:30 in the morning in India and evening in the Americas.

When you hit a limit, the API returns 429 with a Retry-After header. If the daily allowance is spent, it also sends x-should-retry: false, and retrying is pointless until the reset. Response headers such as x-ratelimit-remaining-tokens show how close you are. Reflection says details on requesting higher limits are "coming soon."

Is Reflection AI Beam free? Pricing today and after release

There are two answers, because Beam will exist in two forms.

The weights will be free. Once released under Apache 2.0, anyone can download Beam, run it, fine-tune it and build commercial products on it without paying Reflection anything. "Free" stops at the download button, though. Running a 501-billion-parameter model costs real money in hardware or cloud time, as the next two sections show.

The API has no public price yet. As of October 6, 2026, Reflection has not published per-token rates for Beam. The platform has billing tiers, may ask for a verified card, and can set a daily token allowance, which suggests paid usage is planned, but there is no rate card to quote. If you see a "Beam costs $X per million tokens" figure today, it did not come from Reflection.

For context, here is what the rival models in Reflection's chart cost on their makers' own APIs, per million tokens:

ModelInputOutput
Reflection BeamNot publishedNot published
GLM-5.2 and GLM-5.3 (Z.ai)$1.40 ($0.26 cached)$4.40
Kimi K3 (Moonshot)$3.00 ($0.30 cached)$15.00
DeepSeek V4.1 Flash$0.15 off-peak, $0.30 peak$0.60 off-peak, $1.20 peak

Some early write-ups list GLM-5.3 at a lower price than GLM-5.2. Z.ai's own pricing page lists both at the same $1.40 input and $4.40 output. If Beam's efficiency claim holds, Reflection has room to price below GLM. Whether it does is the most important number still missing.

"So I keep my subscription," Jake said.

"For now, yes," Ethan said. "When Beam has a price, and when independent tests are out, we compare. That's a ten-minute decision next month, not a gamble today."

Can you run Reflection AI Beam locally? The memory math

Short answer: not on a normal PC or laptop, and not on anything yet, because the weights are not out. Once they are, Beam will run on a small class of very large machines. The deciding number is memory, and you can work it out yourself with one rule: memory for the weights ≈ parameters × bytes per parameter. Beam has 501 billion parameters, all of which must be loaded.

PrecisionBits per parameterWeights alone, roughlyQuality
BF16 (full)161,002 GB (933 GiB)Reference
FP88501 GB (467 GiB)Near full; the usual serving format
4-bit (MXFP4, NVFP4, Q4 GGUF)about 4.25 to 4.8about 266 to 301 GB (250 to 280 GiB)Small loss, usually fine for chat and code
3-bit GGUFabout 3.1 to 3.9about 192 to 245 GBNoticeable loss on hard tasks
2-bit GGUFabout 2.1 to 3.0about 130 to 185 GBLarge loss; for experiments only

Three things make the real requirement bigger than this table. Real quantized files usually come out larger than the bare math, because quantizers keep some sensitive layers at higher precision. You need extra memory for the context (the KV cache, which grows with conversation length), and Reflection has not yet published the attention details needed to size it. And the operating system and your other apps need room too. A safe rule is to add 10 to 20% on top of the weights, more for long contexts.

Which machines can actually hold Beam

Your hardwareMemoryCan it run Beam?
Gaming PC or laptop, any GPU16 to 64 GB RAM, 8 to 32 GB VRAMNo. Even 2-bit is several times too big.
128 GB unified memory: Ryzen AI Max+ 395 (Strix Halo), DGX Spark, 128 GB Macs128 GB sharedNo in practice. Only the most extreme 2-bit file would come close, with almost no room for context.
Two linked DGX Sparks, or a 256 GB Mac Studio256 GBMaybe at 3-bit, tightly. Not at 4-bit.
Mac Studio with M5 Ultra, 512 GB512 GB unifiedYes at 4-bit with room for context; FP8 would be too tight.
Workstation with 4× 96 GB GPUs (RTX PRO 6000 class)384 GB VRAMYes at 4-bit, fast.
Server with 8× 80 GB or larger GPUs640 GB and upYes at FP8, the way Reflection-class models are meant to be served.

Apple's Mac Studio with M5 Ultra is offered with up to 512 GB of unified memory, which makes it the simplest single box that can hold a 4-bit Beam. It will not be cheap; even well short of the top memory option, a loaded M5 Ultra runs well into five figures. Everything else that fits is either a multi-GPU workstation or a server.

"Wait," Jake said, "doesn't 23B active mean it runs like a 23B model?"

"For speed, roughly yes," Ethan said. "For memory, no. The speed is set by how much it reads per word. The memory is set by everything it might read."

How fast would it be?

Speed follows the active parameters. For each token, Beam reads its 23 billion active weights from memory. At 4-bit, that is roughly 13 GB per token. A machine's memory bandwidth divided by that figure gives a theoretical ceiling: about 60 tokens per second on hardware with around 800 GB/s, and far more on data-center GPUs with several terabytes per second. Real runs land well below the ceiling once routing, attention and long contexts are added, but this is why a 501B MoE can feel quick on a big Mac, while a dense model of the same size would crawl.

The flip side: if the model does not fit and spills onto an SSD, speed collapses to seconds per token. "It loads from disk" is not a way to run Beam; it is a way to watch a progress bar.

Windows and Kali Linux: what to expect on release day

Right now there is nothing to install on Windows 11 or Kali. When the weights arrive, these are the usual stages, and each one takes time:

  1. Official weights appear on Hugging Face, probably in a full-precision and possibly an FP8 version, with the model card and license.
  2. Server engines such as vLLM and SGLang add support. Reflection promised "integration with a broad range of open source libraries" at launch, so this may happen on day one. These engines run on Linux with NVIDIA GPUs; on Windows, they run inside WSL2.
  3. llama.cpp support follows. Beam's architecture is new, with interleaved local and global attention, its own expert routing and new normalization choices, so llama.cpp needs code changes before any GGUF file works. Expect GGUF quantizations from the usual community publishers once that lands.
  4. Ollama and LM Studio pick it up after llama.cpp. Any "Beam" file you see in those apps before then is either mislabeled or an unrelated model.

On Kali specifically, install vLLM inside a Python virtual environment rather than with system pip; Kali blocks system-wide pip installs, as our guide to the externally-managed-environment error explains. On Windows, use WSL2 with the NVIDIA driver installed on the Windows side.

If you have a 16 to 64 GB machine and want a strong local model today, a smaller open model is the honest answer. Our guide to which big models really run on a PC maps RAM to the models that fit.

How to run Reflection AI Beam on AWS

For most companies, "self-hosting" Beam will mean renting GPUs in the cloud rather than buying a server. AWS is the obvious place to look, and there are three questions to answer: can you use Beam through Amazon Bedrock, which EC2 instance fits it, and what will it cost?

Is Beam on Amazon Bedrock?

No, not as of October 6, 2026. Bedrock is AWS's service for calling ready-made models through an API, without managing any servers, and Beam is not in its model catalog. Reflection says it will launch the weights "with an ecosystem of distribution partners," but it has not named them. If AWS adds Beam to Bedrock later, that will be the easiest route by far, and this page will say so.

You also cannot bring Beam into Bedrock yourself. Bedrock's Custom Model Import feature, which lets you upload your own open-weight model, has three rules Beam breaks today:

  • It accepts only listed architectures: Llama, Mistral, Mixtral, Flan, GPTBigCode, Qwen and GPT-OSS families. Beam's architecture is not on the list.
  • Text models must have weights under 200 GB. Beam is about 501 GB even at FP8.
  • The model's maximum context length must be under 128K. Beam's is 256K on the API and 1M by design.

So for now, Beam on AWS means EC2, the service where you rent virtual machines, including ones with GPUs. If you are new to it, our plain-English explainer on Bedrock versus SageMaker AI explains how the AWS AI services differ.

Which EC2 instance fits Beam, and what it costs

The instance has to hold the weights in GPU memory with room left for context. These are the GPU instances that can, with AWS's on-demand Linux prices in US East (N. Virginia) on October 6, 2026. The monthly column assumes the instance runs around the clock (730 hours).

InstanceGPUsGPU memoryBeam fits atPer hourPer month, 24/7
g7e.24xlarge4× RTX PRO 6000 Blackwell384 GiB4-bit$16.57about $12,100
g6e.48xlarge8× L40S384 GiB4-bit$30.13about $22,000
g7e.48xlarge8× RTX PRO 6000 Blackwell768 GiBFP8, with room$33.14about $24,200
p5.48xlarge8× H100640 GiBFP8$55.04about $40,200
p5en.48xlarge8× H2001,128 GiBFP8 roomy, BF16 tight$63.30about $46,200
p6-b200.48xlarge8× B2001,440 GiBBF16$113.93about $83,200

Some quick reading of that table:

  • The cheapest real option is g7e.24xlarge, at about $16.57 an hour for a 4-bit copy. That is roughly $133 for an eight-hour working day, which makes a one-day evaluation affordable for most teams. It is also cheaper and newer than g6e.48xlarge for the same 384 GiB, so pick g6e only if g7e is not available in your Region.
  • g7e.48xlarge is the sweet spot for FP8. At about $33 an hour it holds the near-full-quality FP8 weights (about 467 GiB) with roughly 300 GiB to spare for long contexts and many users. On paper it beats p5.48xlarge on memory and price; the H100s in P5 have much faster memory and interconnects, which matters when many users share one server.
  • Full precision needs P5en or P6. Unless you are fine-tuning or doing research that needs the full BF16 weights, FP8 is the practical serving format.
  • Running 24/7 is expensive. Even the smallest option is about $12,000 a month if left on. Self-hosting pays off for constant, heavy use, or when data must stay in your own AWS account; for occasional use, an API is almost always cheaper.

Prices differ by Region, and Savings Plans or Reserved Instances lower them for long commitments. Our Reserved vs Savings Plans vs Spot decision tree walks through which discount fits which pattern.

Step by step: launching Beam on EC2 once the weights are out

This is the path we would follow on release day. The model name is a placeholder until Reflection publishes the official repository, and the exact serving flags will come from Reflection's model card; check both before you run anything.

  1. Request GPU quota first. New and small AWS accounts usually have little or no quota for large GPU instances. In the Service Quotas console, find Running On-Demand G and VT instances (for g7e and g6e) or Running On-Demand P instances (for p5 and p6). Quotas are counted in vCPUs: g7e.24xlarge needs 96, and the 48xlarge sizes need 192. Approval can take a day or more, so ask before you need it. Our guide to the vCPU quota exceeded error covers how to word the request.
  2. Plan for capacity. Even with quota, big GPUs can be sold out in a Region, which shows up as InsufficientInstanceCapacity. Try another Availability Zone or Region, or reserve time ahead with EC2 Capacity Blocks for ML, which AWS offers for P-family instances.
  3. Launch the instance with an AWS Deep Learning AMI for Ubuntu, which comes with the NVIDIA driver and CUDA installed. Put it in a private subnet if you can.
  4. Give it storage. Attach a gp3 EBS volume big enough for the weights plus room to spare: about 1 TB for FP8 or a 4-bit copy, about 2 TB if you also want the BF16 weights. At $0.08 per GB-month in US East, 1 TB of gp3 is about $80 a month, and downloading into EC2 from the internet has no data-transfer charge.
  5. Download the weights with the Hugging Face command-line tool:
    pip install -U huggingface_hub
    hf download <official-beam-repo> --local-dir /data/beam
  6. Serve it with vLLM (or SGLang), spreading the model across all GPUs with tensor parallelism:
    python3 -m venv ~/beam-env && source ~/beam-env/bin/activate
    pip install -U vllm
    vllm serve /data/beam --tensor-parallel-size 8 --max-model-len 131072 --port 8000
    Use --tensor-parallel-size 4 on g7e.24xlarge. Start with a shorter --max-model-len and raise it once the model loads; context is what eats the spare memory.
  7. Keep it private. Do not open port 8000 to the internet. Reach it through an SSH tunnel or AWS Systems Manager port forwarding:
    aws ssm start-session --target <your-instance-id> \
      --document-name AWS-StartPortForwardingSession \
      --parameters '{"portNumber":["8000"],"localPortNumber":["8000"]}'
    Then point any OpenAI-compatible tool at http://localhost:8000/v1. Our Systems Manager explainer covers the setup.
  8. Stop it when you are done. A stopped instance does not bill for compute, but its EBS volume still costs storage. Set a budget alert before you start; a forgotten p5 burns about $1,300 a day. Our three-layer billing alert setup takes ten minutes.

Spot Instances look tempting at a fraction of the price, but GPU Spot capacity is scarce, and AWS can take the instance back with a two-minute warning. They are fine for experiments and terrible for anything users depend on.

What about SageMaker AI?

Amazon SageMaker AI can host the same model as a managed endpoint, using AWS's large-model inference containers on the same GPU families. You pay a premium over raw EC2, and in return AWS handles more of the serving, scaling and monitoring. It makes sense for teams already running models in SageMaker. For a first evaluation, plain EC2 is simpler and cheaper.

API or self-host? A simple way to decide

Since Beam has no public API price yet, nobody can give you a break-even number today. But the logic is easy:

  • Use the API if your usage is occasional or bursty, you want to try Beam without commitment, or you do not have someone to run GPU servers.
  • Self-host on AWS if your data must stay inside your own AWS account for legal or customer reasons, you need to fine-tune Beam on your own code or documents, or you will keep a server busy most of the day.
  • Do the math once both numbers exist. Take the instance's monthly cost, divide by the tokens it can serve in a month at your usage pattern, and compare with Reflection's per-token price when it is published.

Beam vs Kimi K3 vs GLM-5.3 vs DeepSeek V4.1 Flash

Beam enters a crowded field of big open-weight models, and most of the strongest ones come from Chinese labs. Here is how they line up on the things that decide what you can actually do with them.

ModelMakerSize (total / active)LicenseDownload today?Images?
BeamReflection AI (US)501B / 23BApache 2.0 (promised)No, later in OctoberNo, text only
Kimi K3Moonshot AI (China)2.8T / 104BCustom Kimi K3 licenseYesYes, and video
GLM-5.3Z.ai (China)about 753B totalCustom GLM-5.3 licenseYesNo
GLM-5.2Z.ai (China)about 753B totalMITYesNo
DeepSeek V4.1 FlashDeepSeek (China)552B backbone / 8B to 16BMITYesNo

What that means in practice:

  • For raw capability today, Kimi K3 leads Reflection's own chart, but at 2.8 trillion parameters it is a data-center model, and its API is the most expensive of the group.
  • For cheap, strong coding today, DeepSeek V4.1 Flash beats Beam on Reflection's coding rows and costs cents per million tokens. Its weights are available under MIT.
  • For a permissive license and a Western supply chain, Beam is the one to watch. Apache 2.0 from a US lab matters to companies and governments whose policies rule out Chinese-origin models, and the smaller total size makes it cheaper to host than GLM-5.3 or Kimi K3.
  • For self-hosting cost, Beam should win among models of its tier: about two-thirds the total size of GLM-5.3, and a fifth of Kimi K3.

Reflection AI vs DeepSeek

The comparison everyone is making. DeepSeek is the Chinese lab whose efficient open models shook the industry, and Reflection's founders have openly framed their company as building a Western answer. On today's evidence, DeepSeek V4.1 Flash is stronger at coding, cheaper to call and already downloadable. Beam's advantages are its Apache 2.0 license, its US origin, and its strength in tool use and instruction following against GLM-5.2. Neither choice is wrong; they solve different problems.

Reflection AI vs Anthropic, OpenAI and Cursor

People search these pairings, so here is the plain difference. Anthropic and OpenAI build closed frontier models that you use only through their apps and APIs; you never get the weights. Reflection is betting on the opposite approach: give the weights away and compete on openness and efficiency. On capability, Reflection's own demos compare Beam with Anthropic's Opus 5 and Fable 5 on a single puzzle, which is not a general ranking. Cursor is a code editor with built-in AI, not a model maker in the same sense, so the realistic relationship is that Beam could one day power tools like it. Today, Beam runs in Reflection's Mirror CLI and in open agents such as OpenCode.

What changes on release day, and what to watch

Reflection says the weights, technical report, model card and developer tools arrive "later this month." When they do, these are the things worth checking, roughly in order of importance:

  1. The license file. Confirm it is plain Apache 2.0 with no extra use policy attached.
  2. The file formats. Is there an official FP8 checkpoint? An official 4-bit one? Official quantizations are usually better than early community ones.
  3. The model card's serving instructions: recommended engine versions, minimum GPU setup, and context settings.
  4. The safety evaluation results, promised with the technical report, plus the safety tests Reflection said it will open-source.
  5. Independent benchmarks. Leaderboards such as Artificial Analysis usually test within days of open weights appearing. That is the moment the efficiency claim gets checked by someone other than its maker.
  6. The API price. The one number that turns "efficient" into "cheap."
  7. Distribution partners. Which clouds and inference providers carry Beam on day one, and whether AWS is one of them.
  8. llama.cpp and GGUF support, which decides when Mac and home-lab users can try it.

We will add a dated update to this page when the weights land, with the real file sizes, repository name and serving commands.

Who should care about Beam, and who can skip it

Watch it closely if you:

  • run AI coding agents and want an open model you can host yourself under a permissive license;
  • work somewhere that cannot use Chinese-origin models but needs open weights for privacy, sovereignty or fine-tuning;
  • pay for large volumes of agent tokens and care about cost per finished task more than peak benchmark scores;
  • build on AWS and want a model that fits on a single 8-GPU instance at FP8.

You can safely skip it, for now, if you:

  • want a model to run on your own laptop or gaming PC; Beam will never fit there, and smaller open models are the answer;
  • need image or document understanding; Beam is text only;
  • are happy with a subscription chatbot for everyday writing; nothing about Beam changes that this month.

Jake landed in the second group, and was relieved to. His real need, faster quotes and listings, was already met. What he took away was a better question to ask about every future headline: "Can I actually download it, and what would it run on?"

Honest limitations, as of October 6

  • No weights yet. Everything about local and AWS hosting is planning, not practice, until the files exist.
  • All benchmarks are self-reported. No independent leaderboard had tested Beam on launch day.
  • No public price. The efficiency claim cannot yet be turned into a cost comparison.
  • Text only. No vision, audio or file input.
  • Context depends on where you read it. 1 million tokens from training, 256K on the beta API, and Reflection notes the API figure "may change during the beta."
  • Behind the top Chinese open models on most coding and reasoning rows in its own chart.
  • Safety results pending until the technical report.
  • Beta API behavior can change, including limits, endpoints and the waitlist.

None of that makes Beam a disappointment. It makes it a version one, from a lab that published its losses next to its wins. That is a better start than most.

Troubleshooting the Reflection API

The errors early users are most likely to hit, and what each one means:

You seeWhat it meansFix
401 invalid_api_keyKey is incomplete, revoked, or its creator left the organizationCreate a new key and copy it in full
401 invalid_authorization_headerHeader is not in the right formSend Authorization: Bearer <key>
403 payment_method_requiredOrganization needs a verified cardAdd one in Billing on the platform
409 billing_setup_required or tier_switch_pendingBilling is incomplete or changing tierFinish setup, or wait for Retry-After
400 unsupported_value on reasoning_effortYou sent none or minimalUse low through max
400 unsupported_parameterA framework sent user, metadata or similarTurn that option off in the framework
400 context_length_exceededPrompt plus output budget is over the limit; nothing is truncated for youShorten history or lower max_completion_tokens
Empty answer, finish_reason: "length"Beam used the whole budget reasoningRaise max_completion_tokens or lower the effort
Output runs past your stop stringstop is accepted but ignoredTrim the text in your own code
415Body sent without a JSON content typeAdd Content-Type: application/json
429 rate_limit_exceededA per-minute or daily limit was hitWait for Retry-After; if x-should-retry: false, wait for 00:00 UTC
503 inference_capacity_unavailableReflection's servers are temporarily fullRetry after Retry-After
"Model not found"Wrong model ID or capitalizationUse exactly Beam-501B-A23B

Every response carries an x-request-id header. Log it; it is the first thing support will ask for.

For IT admins and engineering leads

If your team wants to trial Beam, a little setup now saves an awkward conversation later.

  • Know where the data goes. On the API, prompts and code go to Reflection's servers. Reflection's endpoint does not store completions (store is fixed to false), but review its terms and privacy policy before sending customer data or proprietary source code.
  • Use one project per application and environment. Keys are scoped to a project, which keeps usage and revocation clean. Remember that rate limits are shared across the whole organization, so one runaway agent can starve every other project.
  • Keys belong to people. A key dies when its creator leaves the organization. For production, create keys from a service owner's account and document who that is.
  • Plan the self-hosted path early if your policies require data to stay in your own cloud account. Request GPU quota now, pick a Region with g7e or p5 capacity, and set budget alerts before anyone launches an instance.
  • Pilot with a share of traffic. A new model with self-reported benchmarks belongs on a slice of non-critical work first, with results compared against your current model on your own tasks.
  • Watch the license on release day. Apache 2.0 is easy for legal teams to approve; confirm that is exactly what ships.

Reflection AI the company: stock, valuation and investment questions

A surprising number of people searching for Beam are really asking about the company. The plain facts:

  • Reflection AI is private. It is not listed on any stock exchange, so there is no Reflection AI stock or ticker to buy, and it has not announced plans to go public.
  • Its investors include Nvidia, which led the October 2025 round of $2 billion at an $8 billion valuation, along with firms such as Sequoia Capital and Lightspeed Venture Partners. Its 2026 funding valued it at about $25 billion.
  • Ordinary investors cannot buy in directly. Private-company shares sometimes trade on secondary marketplaces for accredited investors, but that is a niche, risky route, and nothing here is investment advice.
  • How it plans to make money: its API platform, its enterprise and sovereign offerings, and products such as Asimov and Mirror. Giving away weights while selling hosting, support and tools is a common open-model business, but Reflection has not published revenue figures.
  • Is it legit? Yes. It is a well-funded lab founded by senior former DeepMind researchers, with public deals and a real model. Whether Beam is good enough for you is a separate question, and the rest of this page is about that.

Reflection AI Beam: frequently asked questions

What is Reflection AI Beam?

Beam is Reflection AI's first open-weight model, announced October 5, 2026. It is a mixture-of-experts model with 501 billion total and 23 billion active parameters, built for coding, reasoning and agent work, and promised under the Apache 2.0 license.

Is Reflection AI Beam open source?

It will be open-weight under Apache 2.0, which allows commercial use and modification. The training data and most training code stay private, so "open-weight" is the precise term. The weights are not released yet.

When will the Beam weights be released?

Reflection says "later this month," meaning October 2026, together with a technical report, model card and developer tools. It has not given an exact date.

Is Reflection AI free?

The weights will be free to download and use. The beta API has no public price yet and may require a verified payment method. Running the open weights yourself costs money in hardware or cloud GPUs.

How much does the Reflection AI API cost?

Reflection had not published per-token prices as of October 6, 2026. For comparison, GLM-5.3 costs $1.40 per million input tokens and $4.40 per million output tokens on Z.ai's API.

Can I run Reflection AI Beam locally?

Not on a normal PC or laptop. A 4-bit copy needs about 250 to 280 GiB of memory plus room for context, so you need a 512 GB Mac Studio, a multi-GPU workstation or a GPU server, and the weights must be released first.

How much RAM do I need to run Beam?

Roughly 1,002 GB at full BF16 precision, 501 GB at FP8, about 266 to 301 GB at 4-bit, and about 130 to 185 GB at 2-bit, before adding 10 to 20% for context and overhead.

Can I run Beam on AWS?

Yes, on EC2 once the weights are out. A 4-bit copy fits a g7e.24xlarge at about $16.57 an hour in US East, and FP8 fits a g7e.48xlarge at about $33.14 or a p5.48xlarge at $55.04 an hour.

Is Reflection Beam available on Amazon Bedrock?

Not as of October 6, 2026. Bedrock Custom Model Import cannot take it either, because Beam's architecture is not supported, its weights exceed the 200 GB limit for text models, and its context is over the 128K limit.

What does 501B-A23B mean?

501 billion total parameters, with 23 billion active for each token. The active count sets the speed; the total count sets the memory you need.

How does Beam compare to Kimi K3?

Kimi K3 scores higher on almost every shared row in Reflection's chart, handles images and video, and is downloadable today, but it has 2.8 trillion parameters. Beam is much smaller, cheaper to host, and promises a more permissive license.

Is Beam better than DeepSeek?

Not at coding, on Reflection's own numbers: DeepSeek V4.1 Flash scores higher on every shared coding row and costs $0.15 per million input tokens off-peak. Beam's edges are the Apache 2.0 license and its US origin.

Does Beam support images?

No. Beam is text only. The API rejects image, audio and file inputs, though Beam can work with information from other formats once it is turned into text.

What is Beam's context window?

Reflection says training extended it to 1 million tokens, but the beta API allows 256K tokens, counting input and output together, with up to 128K tokens of output.

What is the Beam model ID?

The model ID is Beam-501B-A23B, and it is case-sensitive. Use it with the OpenAI-compatible base URL https://api.reflection.ai/openai/v1.

Is Reflection AI publicly traded?

No. Reflection AI is a private company with no stock listing and no announced plan to go public. Its 2026 funding valued it at about $25 billion.

Who owns Reflection AI?

It was founded by former Google DeepMind researchers Misha Laskin and Ioannis Antonoglou and is backed by investors including Nvidia, which led its $2 billion round in October 2025.

What is Mirror CLI?

Mirror CLI is Reflection's terminal coding agent, in early beta. It uses Beam by default, and you connect it with a Reflection API key.

Can I turn off reasoning in Beam?

No. Beam always reasons. You can choose low, medium, high, xhigh or max effort; none and minimal are rejected with an unsupported_value error.

Is Reflection Beam the same as Reflection 70B?

No. Reflection 70B was a 2024 Llama fine-tune by a different team. Beam is an original model from Reflection AI, the Brooklyn-based lab.

Jake kept his subscription and canceled nothing. A week from now he will know more than any headline told him: whether Beam's weights arrived, what the API costs, and what independent testers found. If a launch like this left you wondering whether you missed something, you didn't. Version one of an open model is a promise and a chart. The download is what makes it yours.

📌 If you keep one line from this page

Active parameters decide how fast a model runs; total parameters decide whether it fits.

Beam is 23B fast and 501B big.

Revision note. Written October 6, 2026, the day after Reflection AI announced Beam. If the headlines had you ready to download it tonight, you were not wrong to be excited; now you know exactly what to wait for.

Related