Meta Muse Glimmer 30B Locally: Ollama, LM Studio, and the 24 GB Truth
Meta Muse Glimmer is the first open model Meta has released since it retired the Llama name, and the first one under a plain Apache 2.0 license. It is a 30-billion-parameter model built for one job: running an AI agent on your own computer, with tool calling, image input and long tasks, no cloud required. The download is about 17 GB, it works in Ollama, LM Studio and llama.cpp on Windows and Kali Linux, and on the right hardware it is the strongest open model of its size for agent work. Here is the surprise the launch coverage skipped: the one-line ollama run muse-glimmer gives you the slow version. The speed Meta advertised, 233 tokens per second on an RTX 5090, needs a second file called the DFlash drafter, and the default Ollama tag does not include it. This guide covers which file to download for your GPU or Mac, the exact steps for each app, how to switch the drafter on, what the model is genuinely good and bad at, and why an 8 GB graphics card is not enough this time.
The evening Jake's shop PC ran Meta's new model at reading speed
Jake runs a phone-repair shop, and the shop PC has a 32 GB of RAM and an 8 GB RTX 4060 that he bought for the point-of-sale software's occasional game of solitaire. When his feed said "Meta's new open model runs on a laptop and beats Google's Gemma at agent tasks," he did what the post told him to do. He typed ollama run muse-glimmer, watched an 18 GB download crawl in over the shop's Wi-Fi, and asked it to read a supplier invoice from a screenshot and draft a reply.
It answered. Correctly, even. It just took most of a minute to produce a paragraph, one word at a time, with the fans at full speed. "Meta said single GPU," he said to Ethan on the phone. "This is a single GPU."
"It said a single GPU," Ethan said. "Not any single GPU. What does Task Manager say about your GPU memory?"
It said 7.8 GB of 8 GB used, and 14 GB of system RAM on top. Ollama had put as many layers as it could on the graphics card and the rest in RAM, where the CPU was doing the work. Ethan explained what Glimmer actually is: a dense 30-billion-parameter model, which means all 30 billion numbers take part in every word it writes, so all of them have to be in fast memory. The models Jake had run before, the 9B MiMo and the small Gemma, fit on the card whole. This one never could.
Two evenings later they ran the same test on Ethan's PC with a 24 GB RTX 3090. The same file answered at 45 tokens a second. Then Ethan switched to the tag with the drafter attached, and the answer came back faster still. Jake's verdict: "So the model was never the problem. The advice was."
This guide is the version of that advice that names the hardware first.
What is Meta Muse? Muse Spark vs Muse Glimmer, and the features that matter
Muse is the name Meta gave its post-Llama model family in April 2026, when Meta Superintelligence Labs replaced Llama inside Meta AI. There are two members you will see named, and the search results mix them up constantly:
| Model | What it is | Open? | How you use it |
|---|---|---|---|
| Muse Spark | Meta's frontier model. Powers Meta AI in Facebook, Instagram and WhatsApp, the Muse Code coding agent and Meta's hosted API. Version 1.3 arrived September 2, 2026 | No | Meta's apps and paid API only |
| Muse Glimmer | A 30B model distilled from Spark, released August 10, 2026 | Yes, Apache 2.0 | Download and run on your own hardware |
So "Meta Muse" the product is Spark, and "Meta Muse" the download is Glimmer. This guide is about Glimmer, the one you can run.
Muse Glimmer comes under the Apache 2.0 license rather than the custom Llama license, so you can use it commercially, modify it and redistribute it with no acceptable-use policy attached. Meta trained it to imitate Spark's outputs (logit distillation during pre-training), then tuned it for a narrow purpose: agents that run locally, keep data on the machine, call tools reliably and recover when a step fails.
Three things make it different from the models you may have run before:
- It sees images. A 1.8B-parameter vision encoder is part of the model, so it reads screenshots, charts, documents and photos. Video is handled as frames (2 per second, up to 96), without sound.
- It is built for agent loops, not chat. Tool calls with strict JSON schemas, long multi-step tasks, and a training emphasis on noticing when something went wrong and retrying.
- Reasoning strength is a dial. You set
low,medium,highorxhighin the system prompt, and the model thinks proportionately.
Its knowledge cutoff is January 4, 2026. It handles 100+ languages, though quality drops outside the well-supported ones. The model card also says it is not intended for users under 18, which matters if you are setting it up in a school or a family PC.
The shock in the spec sheet: 30B, and every parameter is active
Model size numbers have become confusing, and Glimmer is the case where the confusion costs you money. Several popular "30B" models in 2026 are mixture-of-experts designs, written as 30B-A3B: 30 billion parameters exist, but only about 3 billion take part in each token. They still need the whole file in memory, but they run at the speed of a 3B model, which is why they feel fast even from system RAM.
Glimmer is dense. Its 29.6 billion language parameters, plus the 1.8B vision encoder, all work on every token. That has two consequences that decide everything else in this guide:
| Model | Design | Memory the file needs (4-bit) | Speed feels like | Runs from system RAM? |
|---|---|---|---|---|
| Muse Glimmer 30B | Dense 30B | About 17 GB | A 30B model | Only slowly |
| Qwen3.6-27B | Dense 27B | About 16 GB | A 27B model | Only slowly |
| Gemma 4 31B | Dense 31B | About 18 GB | A 31B model | Only slowly |
| Nemotron 3.5 Lightning 30B-A3B | MoE, 3B active | About 17 GB | A 3B model | Yes, usably |
| MiMo-V2.6-Distill-Qwen-9B | Dense 9B | About 6 GB | A 9B model | Yes |
So the question "can I run Muse Glimmer 30B" is really "do I have roughly 20 GB of fast memory in one place." A 24 GB graphics card, a 32 GB Apple Silicon Mac, or one of the new 128 GB unified-memory PCs all say yes. An 8 GB card with 32 GB of RAM says "technically," which is what Jake found.
The upside of dense is quality per gigabyte. Meta's own quantization work shows the 4-bit files losing about 1% on a 15-benchmark average against full precision, and the third-party tests below agree that the 4-bit Glimmer behaves like the full model on agent tasks.
Muse Glimmer 30B hardware requirements: what actually runs it
The system requirements come down to memory. What you need is the model file plus the context cache plus a little overhead, and the Reddit and Hacker News threads are full of people who only checked the first of the three. The file sizes are fixed (next section); the context cache grows with how much text you let the model keep in view. Measured numbers from a Q4_K_M file:
| Machine | Context | Memory used | Generation speed |
|---|---|---|---|
| RTX 3090 (24 GB) | 4K | 16 GB | 45 tokens/s |
| RTX 3090 (24 GB) | 256K | 20 GB | 35 tokens/s |
| RTX 4090 (24 GB) | 4K | about 16 GB | 51 tokens/s |
| RTX 4090 (24 GB), Q4_K_XL, 130K context | 130K | 19.3 GB | 75 tokens/s |
| RTX 5090 (32 GB) | 4K | about 16 GB | 83 tokens/s |
| RTX 5090 (32 GB), with mmproj and DFlash, 64K | 64K | 23.8 GB | 84 tokens/s chat, 121 tokens/s code |
| DGX Spark (GB10, 128 GB unified) | 4K | shared | 13 tokens/s |
| Mac mini M-series, 32 GB, Ollama | default | shared | "slow", usable for walk-away tasks |
Two things stand out. First, Glimmer's sliding-window attention keeps the context cache small: a full 130K-token context on an RTX 4090 fit in 19.3 GB. Most 30B models cannot do that. Second, memory bandwidth decides speed, not GPU compute. The DGX Spark has more memory than any of the cards and is the slowest, because its unified memory is slower than GDDR.
Here is how that maps to hardware you might own:
| You have | Verdict | Which file |
|---|---|---|
| RTX 3090 / 4090 / 7900 XTX (24 GB) | Yes, comfortably | Official Q4_K_M (16.8 GB) or Unsloth UD-Q4_K_XL (15.9 GB) |
| RTX 5090 (32 GB) | Yes, with headroom for vision and DFlash | Official Q4_K_XL (19.7 GB) + mmproj + drafter |
| RTX 4080 / 5080 (16 GB) | Partly on GPU, partly in RAM | Unsloth UD-Q3_K_XL (13.4 GB) or UD-IQ2_M (12.3 GB); expect a big slowdown |
| RTX 4070 / 5070 (12 GB), 3060 (12 GB) | Mostly in RAM | UD-IQ2_XXS (10.7 GB); single-digit tokens per second |
| RTX 4060 / 3050 (8 GB) | Not really | A smaller model. MiMo 9B or Ternary Bonsai 2 fit here |
| Mac, 32 GB unified | Yes | Ollama 30b-mlx (19 GB) or LM Studio 4-bit |
| Mac, 48 GB or more | Yes, full context | Ollama 30b-mlx, or a 6-bit file for better quality |
| Mac, 16 GB | No | The 2-bit file would leave almost nothing for macOS |
| Ryzen AI Max+ 395 laptop or mini PC (64 to 128 GB unified) | Yes | LM Studio with the Vulkan runtime; AMD has a published setup for it |
| CPU only, 32 GB RAM, no GPU | It loads. Do not plan a workflow around it | UD-Q4_K_XL; a few tokens per second at best |
LM Studio's catalog lists 26 GB of RAM as the minimum for its smallest build; Unsloth pitches its 2-bit files at 16 GB cards. Both are honest for what they describe. The 2-bit files run, but quality falls off faster below 3 bits on a dense model than the tidy download size suggests, so an RTX 4080 owner is usually happier on the 3-bit file with some layers in RAM than on the 2-bit file fully on the card.
If you are choosing a new machine for this class of model, our laptop tiers for local LLMs explain why "24 GB of VRAM" is the line that matters.
Muse Glimmer download: which GGUF to get from Hugging Face
There are three sources of GGUF files on Hugging Face, and they are not the same. (The Ollama and LM Studio catalogs pull from these same repositories, so the sizes below are what those apps download too.)
Meta's official repository (meta-models/Muse-Glimmer-30B-GGUF). Two quantized models and the two helper files:
| File | Size | Meant for |
|---|---|---|
| Muse-Glimmer-30B-KQuant-17GB-Q4_K_M.gguf | 16.8 GB | 24 GB cards. The "17GB" quant in Meta's tables |
| Muse-Glimmer-30B-KQuant-Dynamic-Q4_K_XL.gguf | 19.7 GB | 32 GB cards and 48 GB Macs. Slightly better quality |
| mmproj-Muse-Glimmer-30B-Q4_K_M.gguf | 1.4 GB | Vision. Needed for image input in llama.cpp |
| dflash-Muse-Glimmer-30B-Q4_K_M.gguf | 1.63 GB | The speculative-decoding drafter |
Meta's own numbers for these two quants: about 1% average loss across 15 benchmarks versus full precision.
Unsloth's repository (unsloth/Muse-Glimmer-30B-GGUF). A full ladder, built with Unsloth's dynamic method that keeps sensitive layers at higher precision:
| File | Size | Fits on |
|---|---|---|
| UD-IQ2_XXS | 10.7 GB | 12 GB cards, with context in RAM |
| UD-IQ2_XS | 11.5 GB | 12 GB cards, tight |
| UD-IQ2_M | 12.3 GB | 16 GB cards |
| UD-Q2_K_XL | 12.4 GB | 16 GB cards |
| UD-IQ3_XXS | 13.1 GB | 16 GB cards |
| UD-Q3_K_XL | 13.4 GB | 16 GB cards, best 16 GB choice |
| UD-IQ3_M | 14.1 GB | 16 GB cards, some offload |
| UD-Q4_K_XL | 15.9 GB | 24 GB cards. The default recommendation |
| UD-Q5_K_M | 19.2 GB | 24 GB cards with short context, or 32 GB |
| UD-Q5_K_XL | 21.8 GB | 32 GB cards |
| UD-Q6_K_XL | 26.3 GB | 32 GB cards, 48 GB Macs |
| Q8_0 | 29.6 GB | 48 GB and up |
| UD-Q8_K_XL | 32.3 GB | 64 GB and up |
| mmproj-BF16 / mmproj-Q8_0 | 3.85 GB / 2.05 GB | Vision, higher precision than Meta's 1.4 GB file |
LM Studio's community repository carries the official 16.8 GB Q4_K_M and its mmproj, which is what the LM Studio search result downloads.
Which to pick, in one line each: 24 GB card, UD-Q4_K_XL or the official Q4_K_M; 32 GB card, the official Q4_K_XL plus the two helpers; 16 GB card, UD-Q3_K_XL; 12 GB card, UD-IQ2_XXS and low expectations; Mac 48 GB, UD-Q6_K_XL for the best quality that still leaves room for the system.
Download only what you need. The model file, and the mmproj only if you want images, and the DFlash drafter only if your app supports it (llama.cpp does; LM Studio and Ollama handle it through their own tags and runtimes). On a 32 GB card the Q4_K_XL, mmproj and drafter together used 23.8 GB at 64K context in one measured setup, so the helper files are not free.
How to run Muse Glimmer in LM Studio on Windows
LM Studio had launch-day support and the model is in its catalog, so this is the shortest path on Windows.
- Install LM Studio from lmstudio.ai and update if you already have it. Glimmer needs a 2026 runtime; an old install will show "unsupported architecture" until you update the runtime under Settings.
- Open the search (the magnifying glass, or Explore in newer builds) and type Muse Glimmer. Pick the entry from Meta or the lmstudio-community listing.
- Choose the 4-bit file (about 17 GB). LM Studio marks files it thinks will not fit; on a 24 GB card the 4-bit shows as a full fit, on 16 GB it warns.
- Download, then load it from the chat screen. In the load dialog, set the context length. 16K is plenty for chat; agent work wants 64K or more, which costs memory (see the table above).
- Under GPU settings, leave offload at maximum. If the model does not fit, LM Studio will offload as many layers as it can and run the rest on the CPU, which is Jake's walking-pace mode.
- Set the sampling values Meta recommends: temperature 1.0, top-p 0.95, top-k 64. Higher temperature than most models like, and it matters for the reasoning behavior.
- To pick reasoning strength, put
Reasoning strength: highin the system prompt. Uselowfor quick replies andxhighfor coding or multi-step tasks.
For images, drag a screenshot into the chat. LM Studio downloads the vision projector alongside the model, so it works without extra steps. For tools, LM Studio's local server (Developer tab) exposes an OpenAI-compatible API on port 1234 that agent apps can point at; Glimmer's tool calls come through in standard function-call JSON.
LM Studio's own agent test, BionicBench v0.1, is worth knowing because it is one of the few independent-ish comparisons at this size: 18 tasks covering repository instructions, attachments, screenshots, editing Word, PowerPoint and Excel files, and generating PDFs. Glimmer completed 83.3%; Gemma 4 31B and Qwen 3.6 27B each completed 77.7%. It is one company's test of 18 tasks, so treat it as a signal, not a verdict.
How to run Muse Glimmer with Ollama on Windows and Kali
Ollama has an official library entry, which is rarer than it sounds for a new model (Bonsai 2 and MiMo 9B are still community-only). The catch is that there are 15 tags, and the default is not the best one.
The steps, on either system:
- Install Ollama. Windows: the installer from ollama.com. Kali:
curl -fsSL https://ollama.com/install.sh | sh, which detects an NVIDIA or AMD GPU and installs the right runtime. - Check
ollama --version. You want a 2026 build; older ones do not know the architecture. - Pick a tag from the table below by your memory, and pull it with
ollama run <tag>. The first run downloads 18 to 20 GB. - Raise the context and set reasoning strength (commands below), then check
ollama psto confirm100% GPU. - Point your agent app at
http://localhost:11434/v1with the tag as the model name.
Pick the tag. All tags are 128K context.
| Tag | Size | What it is | Use it when |
|---|---|---|---|
muse-glimmer / :30b / :30b-q4_K_M | 18 GB | 4-bit, no drafter | Default. Fine, slower than it could be |
:30b-q4_K_M-dflash | 20 GB | 4-bit with the DFlash drafter | You have 24 GB or more and want speed |
:30b-q8_0 | 31 GB | 8-bit | 48 GB Macs, dual-GPU rigs |
:30b-q8_0-dflash | 33 GB | 8-bit with drafter | Same, faster |
:30b-bf16 / :30b-bf16-dflash | 57 / 59 GB | Full precision | 64 GB and up; testing, not daily use |
:30b-mlx | 19 GB | Apple MLX 4-bit | Apple Silicon Macs (fastest path on a Mac) |
:30b-nvfp4 / -dflash | 17 / 19 GB | MLX NVFP4 | Apple Silicon, slightly smaller |
:30b-mxfp8 / -dflash | 33 / 35 GB | MLX 8-bit | 64 GB and up Macs |
:30b-mlx-bf16 / -dflash | 60 / 65 GB | MLX full precision | 128 GB Macs |
Run it:
ollama run muse-glimmer:30b-q4_K_M-dflash
Ollama pulls the 20 GB, loads it, and gives you a prompt. Type /set parameter num_ctx 32768 to raise the context from Ollama's default, and /set system "Reasoning strength: high" to set the thinking level for the session. To make those permanent, write a Modelfile:
FROM muse-glimmer:30b-q4_K_M-dflash
PARAMETER num_ctx 32768
PARAMETER temperature 1.0
PARAMETER top_p 0.95
PARAMETER top_k 64
SYSTEM "Reasoning strength: high"
and ollama create glimmer-agent -f Modelfile. Now ollama run glimmer-agent has your settings every time, and any app that talks to Ollama's API can use glimmer-agent as the model name.
Images: in ollama run, paste a file path on its own line after your question (Windows: C:\Users\Jake\Desktop\invoice.png) and Ollama sends it as an image. Through the API, images go in the images array as base64.
Check where it is running. ollama ps shows the loaded model and a column like 100% GPU or 62%/38% CPU/GPU. Anything but 100% GPU is the split mode, and that is why it is slow. There is no setting that fixes a memory shortfall; the fix is a smaller file or more memory.
🙋♂️ Jake's Reality Check
Jake's shop PC showed 44%/56% CPU/GPU for the 18 GB default tag. Ethan's RTX 3090 showed 100% GPU for the 20 GB dflash tag with 16K context, and dropped to a CPU/GPU split when Jake pushed context to 128K "because the tag said it could." It can. Whether the card can hold the cache at the same time is a separate question, and ollama ps is where you find out.
Running Muse Glimmer with llama.cpp (Windows and Kali), with vision and the drafter
llama.cpp is where the DFlash drafter, the vision projector and the exact memory settings are all under your control. It is also where the launch-day error lives.
Get a new enough build. Support for the muse-glimmer architecture was merged on August 10, 2026 (pull request #26841). Builds from b10344 and earlier fail with:
llama_model_load: error loading model: unknown model architecture: 'muse-glimmer'
Anything from b10362 onward loads it. On Windows, download the release zip for your GPU (CUDA for NVIDIA, Vulkan for AMD or Intel) from the llama.cpp releases page. On Kali, either take the Vulkan or CUDA release zip, or build:
git clone https://github.com/ggml-org/llama.cpp
cmake llama.cpp -B llama.cpp/build -DGGML_CUDA=ON
cmake --build llama.cpp/build --config Release -j
Swap -DGGML_CUDA=ON for -DGGML_VULKAN=ON on AMD, or drop the flag for CPU-only.
The one-liner that pulls and serves. Newer llama.cpp builds fetch from Hugging Face directly:
llama-server -hf meta-models/Muse-Glimmer-30B-GGUF --port 8080
That downloads the official Q4_K_M and its mmproj, and opens a web chat at http://localhost:8080 plus an OpenAI-compatible API at /v1. On Kali the same command works once the binary is on your PATH.
With vision and the drafter, explicitly:
llama-server -m Muse-Glimmer-30B-KQuant-17GB-Q4_K_M.gguf ^
--mmproj mmproj-Muse-Glimmer-30B-Q4_K_M.gguf ^
--spec-type draft-dflash --spec-draft-n-max 15 ^
-c 32768 -ngl 99 --temp 1.0 --top-p 0.95 --top-k 64 --port 8080
(Use \ instead of ^ for line continuation on Kali.) -ngl 99 puts every layer on the GPU; lower it to leave some in RAM on a 16 GB card. --spec-draft-n-max 15 is the maximum for DFlash's block size of 16 (one anchor token plus 15 proposals); larger values are clamped.
Unsloth's files work the same way with -hf unsloth/Muse-Glimmer-30B-GGUF:UD-Q4_K_XL, and their higher-precision mmproj-BF16.gguf gives slightly better document reading at the cost of 2.5 GB more memory.
Vulkan on Kali and the drafter. The DFlash path was written and tested first on CUDA and Metal. If speculative decoding makes output slower or odd on a Vulkan build, run without the --spec-* flags. You lose the speed-up, not the model.
Why Ollama's default tag is slower than Meta's numbers: DFlash explained
Meta's headline speeds are 233 tokens per second on an RTX 5090, 50 on an M5 Max and 38 on an M4 Max. The same hardware without the trick does 75, 27 and 24. The trick is DFlash, a speculative-decoding drafter that ships as a separate 1.63 GB file.
Speculative decoding works like an autocomplete that guesses ahead. The small drafter proposes a block of up to 15 tokens, and the big model checks the whole block in one pass, which is far cheaper than producing the tokens one at a time. Every accepted token is free speed; a rejected one costs a little. Meta measured a 3.1x gain on the RTX 5090 and 1.5x to 1.8x on Macs; an independent test on an RTX 5090 with a 64K context saw 84 tokens per second in conversation and 121 in code, where the drafter guesses well because code is predictable.
The output is identical either way. Speculative decoding does not change what the model would have said; it only changes how fast it says it.
What this means per app:
- Ollama: use a
-dflashtag. The plain tag is the model alone. - llama.cpp: add
--spec-type draft-dflash --spec-draft-n-max 15and have the drafter file present (the-hffetch of the official repository includes it). - LM Studio: the runtime handles speculative decoding in its own way; check the model's load options for a draft-model setting in your version.
- Any app on an 8 GB or 12 GB card: skip it. The drafter takes memory you do not have, and speculative decoding does not help when the bottleneck is the CPU/GPU split.
Reasoning strength, context window and the settings that matter
Glimmer's controls are simpler than most 2026 models, and a couple of them are unusual.
Reasoning strength is set with a line in the system prompt, exactly Reasoning strength: low (or medium, high, xhigh). There is no separate thinking toggle. Meta suggests high or xhigh for coding and agent tasks. Low is fast and fine for summaries, rewrites and simple questions; xhigh can spend thousands of tokens thinking before it answers, which is where a 45-tokens-per-second machine starts to feel slow.
Sampling should be temperature 1.0, top-p 0.95, top-k 64. That temperature is higher than the 0.6 or 0.7 many people use by habit. Lowering it makes Glimmer more repetitive in long agent loops rather than more accurate.
Context window is 131,072 tokens by default in the official files, and Unsloth's build allows up to 262,144. You rarely need either; 32K covers a long coding session, and each token of context costs memory. The sliding-window design keeps that cost lower than usual, which is why 130K fit on a 24 GB card in one test, but set what you use.
In Transformers, the same dial is a parameter on the chat template call: reasoning_strength="low" inside apply_chat_template. The chat template also accepts a tools list in the OpenAI function format.
Using Muse Glimmer as a local agent: tool calling, OpenClaw, Hermes and OpenCode
This is the job Glimmer was built for, and the reason to pick it over a general chat model of the same size.
Tool calling uses standard function definitions: a tools list with type: function, a name, a description and a JSON-schema parameters block. The model returns a tool call as JSON, you run the tool, return the result as a tool message, and it continues. Reviewers who put it into real agent frameworks report the specific thing you want at this size: it does not mangle the call format, does not invent tools that were not offered, and picks the right tool at the right time. That reliability is the whole game in a local agent, because a 30B model that calls the wrong tool 5% of the time fails a 20-step task most of the time.
Local API. All three apps expose an OpenAI-style endpoint (Ollama on 11434, LM Studio on 1234, llama.cpp on 8080). Any agent tool that lets you set a base URL and a model name will work: point it at http://localhost:11434/v1, model glimmer-agent (your Modelfile name) or muse-glimmer:30b-q4_K_M-dflash.
OpenClaw was one of the two scaffolds Meta named at launch. Its config file is ~/.openclaw/openclaw.json (Windows: under your user profile). Add a provider entry with your local base URL and model, then openclaw gateway restart. Hermes Agent is the other, and works the same way through its provider settings. OpenCode is the coding-agent case most tested in public write-ups: pointed at a local llama.cpp server, Glimmer handles multi-file edits and repository instructions, with one consistent note that it does best when the task is spelled out step by step rather than left open.
What a local agent can do with it. Read the PDF in your Downloads folder and file it; watch a folder and rename screenshots by their content; run a build, read the error, fix the file, run again; fill a spreadsheet from a pile of invoices. All without the data leaving the machine. That is the "always-on local agent" Meta describes, and the privacy angle is real: the model card's training notes emphasize data minimization and local-first execution.
What to be careful about. Meta's own guidance is that Glimmer should run inside a system with guardrails, not as a bare endpoint, and that irreversible actions (deleting, sending, paying) should need a human click. Its prompt-injection resistance scored better than Qwen3.6-27B and worse than Gemma 4 31B on the AgentDojo test (attack success 28.4% versus 40.3% and 25.6%), and on the CI Memories privacy test it leaked more than Gemma. A local agent that reads untrusted web pages or email is exactly the scenario those numbers describe. Keep it on a limited account and review what it does. If you want the plain-English version of what an agent is and why these tests exist, our honest answer on AI agents covers it.
Muse Glimmer with images: screenshots, documents and charts
The vision encoder is a 1.8B ViT-G/14, 50 layers, and it is genuinely good for its size: 78.8 on CharXiv reasoning (charts), 75.8 on OmniDocBench (documents), 75.4 on ScreenSpot Pro (finding things on screenshots), 74 on MMMU Pro. All three of the 30B-class open models cluster within a few points here, so pick on other grounds.
Practical notes:
- Each image can use up to 4,096 visual tokens, so a page of dense text costs as much context as a few thousand words. Budget for it.
- Video is frames only, 2 per second up to 96 frames (48 seconds), no audio.
- Screenshot reading is the agent use case that works best: "what does this error dialog say and what should I click" is a solved problem at this size.
- Object detection returns coordinates in Meta's native format; the exact schema is in the model card examples rather than documented separately, so test it before building on it.
In llama.cpp, vision needs the --mmproj file. In Ollama and LM Studio it is built into the tag. In all three, if images silently do nothing, the projector was not loaded.
Muse Glimmer benchmarks: honest strengths and weak spots
Meta compares Glimmer with the two dense open models of the same size, Gemma 4 31B and Qwen3.6-27B. Here is the full picture, including where it loses.
| Benchmark | Muse Glimmer 30B | Gemma 4 31B | Qwen3.6-27B |
|---|---|---|---|
| MCP Atlas (agent tool use) | 75.5 | 54.2 | 62.5 |
| DeepSearch QA | 74.6 | 61.7 | 71.1 |
| WildClawBench | 47.6 | 37.6 | 43.2 |
| Gaia2 | 43.3 | 36.4 | 40.0 |
| SWE-Bench Pro | 51.2 | 36.9 | 50.2 |
| SWE-Bench Verified | 76.0 | 66.6 | 77.2 |
| Terminal-Bench 2.1 | 51.7 | 43.4 | 60.7 |
| OSWorld-Verified (computer use) | 65.9 | 58.5 | 75.6 |
| SkillsBench | 44.3 | 32.4 | 46.6 |
| GDPVal-AA v2 | 953 | 811 | 1141 |
| SciCode | 43.6 | 43.4 | 39.8 |
| IFBench (instruction following) | 77.0 | 76.0 | 70.8 |
| AIME 2026 (math) | 94.7 | 89.2 | 94.1 |
| GPQA Diamond | 83.5 | 85.7 | 84.2 |
| HLE text | 22.0 | 23.6 | 23.1 |
| AA-LCR (long context) | 80.0 | 68.3 | 73.3 |
| CharXiv reasoning (vision) | 78.8 | 77.7 | 78.4 |
| AgentDojo attack success (lower is better) | 28.4 | 25.6 | 40.3 |
| CI Memories privacy violation (lower is better) | 26.4 | 12.1 | 53.4 |
Read it as three stories.
Where Glimmer clearly wins: tool use and search-style agent tasks (MCP Atlas by 13 points over Qwen, DeepSearch QA), instruction following, long context, and math. This is the training target, and the numbers show it.
Where Qwen3.6-27B is better: terminal work (Terminal-Bench 2.1, 60.7 versus 51.7) and computer use (OSWorld, 75.6 versus 65.9). If your agent lives in a shell, Qwen is still the stronger pick at this size, and a Hacker News commenter made exactly that point on launch day.
Where Gemma 4 31B is better: safety and privacy tests, and a little on graduate-level science questions. Gemma leaks less and resists injection slightly better.
Two more data points from outside Meta: LM Studio's 18-task BionicBench put Glimmer at 83.3% against 77.7% for both rivals, and independent quantization testing shows the 4-bit file within about 1% of full precision, so the numbers above survive the download.
Muse Glimmer on Kali Linux: the security angle, and the VM problem
Kali readers get the same Ollama and llama.cpp steps as above, with two Kali-specific realities.
The VM problem. Most Kali installs are virtual machines, and a VM does not see the host's GPU unless you set up passthrough, which is rare on laptops. Inside a VM, Glimmer runs CPU-only: it loads if you gave the VM 24 GB or more of RAM, and it produces a few tokens per second. That is enough to test a prompt, not enough to run an agent. For serious use, run Ollama on the host (Windows or a bare-metal Linux) and point tools inside the VM at the host's IP on port 11434, with OLLAMA_HOST=0.0.0.0 set on the host.
Bare-metal Kali with an NVIDIA card is the good case. Install the NVIDIA driver from Kali's repository (sudo apt install nvidia-driver nvidia-cuda-toolkit), reboot, confirm nvidia-smi sees the card, then the Ollama install script picks up CUDA automatically. AMD cards go through Vulkan or ROCm; the Vulkan release of llama.cpp is the least fuss.
What it is useful for on Kali. Explaining tool output (an nmap scan, a Burp request, a stack trace), drafting the report section for a finding, reading code for a review, turning a screenshot of a dialog into text. The agent side, given a shell tool, can run a scan, read the result and run the next step, which is exactly where the guardrail advice above applies twice: a local agent with a shell on a security box needs a human approving each command. Glimmer has no security-specific training that we know of; MiMo 9B was trained deliberately on security data and is the better small model for that, while Glimmer is the better general agent.
As always, use any of this only on systems you are authorized to test.
Muse Glimmer on a Mac
Apple Silicon is the other hardware Meta named at launch, and it is the smoothest experience below a 24 GB PC card, because unified memory means a 32 GB MacBook has "32 GB of VRAM" for this purpose.
- Ollama:
ollama run muse-glimmer:30b-mlx(19 GB). The MLX tags run on Apple's framework and are faster than the GGUF tags on the same Mac. The-dflashMLX variants add the drafter. - LM Studio: pick the MLX build in the download list rather than the GGUF one.
- Speeds Meta measured: M4 Max 24 tokens per second, 38 with DFlash; M5 Max 27, 50 with DFlash. An M-series Mac with 32 GB will be slower than a Max chip because memory bandwidth scales with the chip tier.
- Memory: 32 GB is the floor for the 4-bit file with a working context. 48 GB lets you run a 6-bit file or a long context. 16 GB Macs should skip this model.
One HN reader's description of a 32 GB Mac mini on the default tag was that everything runs, slowly enough to "go for a walk" while it works. That is the honest Mac mini experience, and it is still a private, offline agent for the price of patience.
Muse Glimmer pricing: hosted API, vLLM on a server, and when the cloud makes sense
If your machine fails the hardware table, you can still use Glimmer without buying anything. Meta lists Fireworks, Together AI and OpenRouter as hosted providers, all with OpenAI-compatible APIs, and all much faster than a home GPU. Prices change monthly; check the provider's page rather than any number in a blog post. The Ollama library has no cloud tag for this model, so Ollama's own cloud service is not an option for it at the time of writing.
If you have a server with a real GPU and want to serve a team, vLLM is the route Meta documents: vllm serve meta-models/Muse-Glimmer-30B --model-impl transformers (add --tensor-parallel-size for multiple cards), which loads the full-precision weights, so budget about 60 GB of GPU memory plus context. SGLang and NVIDIA's NIM container are the other two server paths Meta and NVIDIA published, and on Blackwell data-center cards the model reaches over 20,000 tokens per second per GPU in batched serving. None of that is a home setup; it is the reason a hosted Glimmer costs so little per token.
The trade is the obvious one. Hosted means the data leaves your machine, which removes the main reason to want Glimmer over a larger cloud model. Hosted Glimmer makes sense for testing whether the model suits your task before you buy a card, or for a workload where the input is not sensitive and you want the cheapest capable agent model. Local makes sense for everything Meta built it for.
There is no paid Meta API for Glimmer specifically; Meta's own hosted offering is Muse Spark, the closed teacher model, which is a different product with its own pricing.
Muse Glimmer not working? Fixes for the common errors
| What you see | Cause | Fix |
|---|---|---|
unknown model architecture: 'muse-glimmer' | llama.cpp older than b10362 (or an app built on it) | Update llama.cpp; in LM Studio update the runtime; in Jan or other wrappers wait for the bundled llama.cpp to move |
| Model loads, runs at 1 to 3 tokens per second | Split across GPU and RAM; not enough VRAM | ollama ps or the LM Studio load screen shows the split. Use a smaller file, shorter context, or accept it |
ollama run hangs at the download, or "pull model manifest" errors | 18 to 20 GB pull over a flaky connection | Re-run; Ollama resumes. Check free disk; Ollama needs the full size plus temp space |
| Out of memory as soon as you send a long prompt | Context cache pushed total memory over the card | Lower num_ctx (Ollama), -c (llama.cpp) or context length (LM Studio). 16K to 32K is realistic on 24 GB |
| Images ignored, model describes "the text you sent" | Projector not loaded | llama.cpp: add --mmproj. Ollama/LM Studio: use an official tag, not a custom import |
| Thinking goes on for minutes | xhigh reasoning on a slow machine | Set Reasoning strength: medium or low in the system prompt |
| Tool calls come back as plain text instead of JSON | App not passing the tools list through, or custom template | Use the official tag or the official GGUF's template; test with a curl call to the local /v1/chat/completions endpoint with tools set |
| Output gets repetitive in long tasks | Temperature lowered by habit | Put it back to 1.0, top-p 0.95, top-k 64 |
| DFlash makes it slower or garbles on AMD/Intel | Speculative path on Vulkan | Run without --spec-* flags, or use the non-dflash Ollama tag |
| LM Studio says the model is "unsupported" | Old runtime | Settings, Runtime, update the llama.cpp or MLX runtime, then reload |
For IT admins: Glimmer in a business, and what the license allows
Apache 2.0 is the part that changes the conversation with legal. The Llama licenses had an acceptable-use policy, a 700-million-user clause and attribution rules; Glimmer has none of that. You can run it, fine-tune it, embed it in a product and never tell Meta.
Practical points for a deployment:
- One workstation, one agent. The dense design means a 30B model per machine, so this is not a "run it on every laptop" model. Put it on the machines with 24 GB cards, or on one shared box with Ollama's API exposed on the LAN, and let thin clients talk to it.
- A shared server with a 32 GB or 48 GB card serves a small team through the Ollama or llama.cpp API with
OLLAMA_HOST=0.0.0.0and a firewall rule limiting the port to your subnet. Ollama has no authentication of its own; put it behind a reverse proxy with a key if it leaves the LAN. - Guardrails are on you. Meta's guidance is explicit that Glimmer should run with confirmation for irreversible actions, output filtering where the audience needs it, and your own evaluation set for your use case. The AgentDojo and CI Memories scores above are the reason.
- Age. The model card states it is not intended for under-18 use. Schools and family-facing deployments should note that.
- Data. Everything stays on the machine, which is the point. Log the agent's actions anyway; a local agent with file access is an insider with no memory of what it did.
- Fine-tuning is realistic. LoRA on a single 80 GB card is documented for both supervised tuning and GRPO; full fine-tuning wants eight of them.
Muse Glimmer vs Gemma 4 31B vs Qwen3.6-27B vs the small models: which to install
| You want | Install | Why |
|---|---|---|
| The best local agent on a 24 GB card | Muse Glimmer | Tool-use and search benchmarks, reliable call format, Apache 2.0 |
| A local agent that lives in a terminal | Qwen3.6-27B | Terminal-Bench 60.7 vs 51.7; computer-use lead |
| The safest model to point at untrusted input | Gemma 4 31B | Best injection resistance and privacy scores of the three |
| A big model on a 16 GB card | Ternary Bonsai 2 27B (2-bit native, 6 GB) | Different technology; needs a patched llama.cpp |
| Something for an 8 GB card | MiMo-V2.6 9B or MiniCPM5 2B | Sized for the hardware; MiMo has security training |
| A fast MoE that runs from RAM | Nemotron 3.5 Lightning 30B-A3B | 3B active parameters, so it is quick even when offloaded |
| General chat, not agents | Gemma 4 or Qwen3.6 | Glimmer's training leans toward tasks; the others are rounder conversationalists |
Muse Glimmer vs Qwen 3.8. Meta's comparison table uses Qwen3.6-27B, but the model most people are actually choosing against is the newer Qwen3.8-27B (August 2026), which has an official Ollama tag and is the dense 27B we covered on Windows and Kali. There is no head-to-head from either company. What can be said honestly: both need the same class of memory (Qwen3.8-27B's 4-bit is about 18 GB), Qwen3.8 is the stronger general-purpose and terminal model on its own card, and Glimmer's edge is specifically tool use, search-style agent tasks and the reliability of its calls. If you can only keep one 18 GB model on the disk, choose by which of those two jobs you do more.
The one-paragraph version: if you have the memory, Glimmer is the agent model to try first in 2026; if you do not, no amount of settings will make a dense 30B comfortable on an 8 GB card, and the smaller models in our run-AI-locally series exist for exactly that machine.
Muse Glimmer review and verdict: should you install it?
Yes, if you have a 24 GB card, a 32 GB Mac or a unified-memory PC and want a local agent. It is the most reliable tool-caller at this size, it reads screenshots, the license is clean, and with the drafter it is fast enough to feel like a service.
Yes, if you are choosing between the three dense 30B models and your work is tool use, research-style tasks or document handling. Glimmer wins those on Meta's tables and LM Studio's.
Maybe, if your agent mostly runs shell commands. Qwen3.6-27B is still ahead there, and the gap is not small.
No, if you have an 8 GB or 12 GB card and expected "runs on a laptop" to mean your laptop. It runs, at a pace that makes a 20-step agent task an afternoon. Pick a model sized for the card, and come back to Glimmer when the hardware changes.
Jake's shop PC went back to MiMo 9B, which fits the card. Ethan's RTX 3090 now runs glimmer-agent with the drafter, reasoning on high, watching a shared folder for supplier PDFs and filing them by vendor and month. "It's slower than the cloud one," Jake said. "And it's never seen a single invoice leave the building," Ethan said. That trade is the entire pitch, and for the first time it comes with a license nobody has to read twice.
Meta Muse Glimmer FAQ
What is Muse Glimmer?
Muse Glimmer is Meta's 30-billion-parameter open-weight model, released August 10, 2026 under the Apache 2.0 license. It is distilled from Meta's closed Muse Spark model and built for local AI agents: tool calling, multi-step tasks, image input and failure recovery, running on a single consumer GPU or a Mac.
Is Muse Glimmer free?
Yes. The weights are free to download from Hugging Face, Ollama and LM Studio, and Apache 2.0 allows commercial use, modification and redistribution with no usage policy. Hosted versions on Fireworks, Together AI and OpenRouter are paid per token.
Is Muse Glimmer open source?
Its weights are open under Apache 2.0, which is more permissive than any Llama license. The training data and the teacher model, Muse Spark, are not released, so it is open-weight rather than fully open source in the strict sense.
Can I run Muse Glimmer 30B locally?
Yes, with about 20 GB of fast memory in one place: a 24 GB graphics card, a Mac with 32 GB of unified memory, or a unified-memory PC. On 16 GB cards it runs partly from system RAM and slows down sharply. On 8 GB cards it loads but is impractical.
How much VRAM does Muse Glimmer need?
The 4-bit file is 16 to 17 GB, and measured use was 16 GB at a 4K context and 20 GB at 256K on an RTX 3090. Plan on 24 GB for comfort, 32 GB if you add vision and the DFlash drafter with a long context.
What is the best GGUF for Muse Glimmer?
For a 24 GB card, Unsloth's UD-Q4_K_XL (15.9 GB) or Meta's official Q4_K_M (16.8 GB). For 32 GB, Meta's Q4_K_XL (19.7 GB). For 16 GB, UD-Q3_K_XL (13.4 GB). Meta's 4-bit files lose about 1% on a 15-benchmark average against full precision.
How do I run Muse Glimmer with Ollama?
Install Ollama, then run ollama run muse-glimmer:30b-q4_K_M-dflash for the 4-bit model with the speed drafter, or ollama run muse-glimmer for the plain 18 GB default. Use /set parameter num_ctx to raise the context, and put "Reasoning strength: high" in the system prompt for agent work.
Does Muse Glimmer work in LM Studio?
Yes, with launch-day support. Search "Muse Glimmer" in LM Studio, download the 4-bit build, and update the runtime if an older install reports an unsupported architecture. LM Studio's own 18-task agent test scored it 83.3% against 77.7% for Gemma 4 31B and Qwen 3.6 27B.
What is DFlash in Muse Glimmer?
DFlash is a 1.63 GB speculative-decoding drafter released alongside the model. It proposes blocks of up to 15 tokens that the main model verifies in one pass, giving Meta's measured 3.1x speed-up on an RTX 5090 and 1.5x to 1.8x on Apple Silicon, with identical output. Ollama offers it as the -dflash tags.
Why is Muse Glimmer slow on my PC?
Almost always because the model does not fit in GPU memory and is split with system RAM. Run ollama ps; anything other than 100% GPU is the split. It is a dense 30B model, so the fix is a smaller quantization, a shorter context or more memory, not a setting.
Is Muse Glimmer better than Gemma 4 or Qwen3.6?
It leads on agent tool use, search-style tasks, instruction following, long context and math. Qwen3.6-27B is better at terminal work and computer use, and Gemma 4 31B scores better on prompt-injection and privacy tests. Pick by the job.
Does Muse Glimmer support images?
Yes. A built-in 1.8B vision encoder reads screenshots, documents, charts and photos, up to 4,096 visual tokens per image. Video is processed as frames, 2 per second up to 96, without audio. In llama.cpp you must load the mmproj file; Ollama and LM Studio include it.
Can Muse Glimmer call tools?
Yes, with standard OpenAI-style function definitions, and reliable formatting is one of its strengths. It works with OpenClaw, Hermes Agent and OpenCode through the local API of Ollama, LM Studio or llama.cpp.
What is Meta Muse? Is Muse Spark the same as Muse Glimmer?
Muse is Meta's post-Llama model family. Muse Spark is the closed frontier model that powers Meta AI and Meta's hosted API; version 1.3 arrived on September 2, 2026 and is not downloadable. Muse Glimmer is the open 30B model distilled from Spark. When people say "Meta Muse" they usually mean Spark; when they download something, it is Glimmer.
Is Muse Glimmer a mixture-of-experts (MoE) model?
No. It is a dense model: all 29.6 billion language parameters, plus the 1.8 billion vision parameters, work on every token. That is why it needs about 20 GB of fast memory and runs slowly from system RAM, unlike 30B-A3B MoE models that only activate 3 billion parameters at a time.
Can Muse Glimmer generate images?
No. It reads images (screenshots, documents, charts, photos) and video frames, and it writes text. It does not produce images, audio or video. For image generation you need a separate model such as Qwen-Image.
Muse Glimmer vs Qwen 3.8: which is better?
There is no official head-to-head; Meta's tables compare with the older Qwen3.6-27B. Both need a 24 GB card at 4-bit. Qwen3.8-27B is the stronger general-purpose and terminal model; Glimmer leads on tool use, search-style agent tasks and reliable function calls. Choose by the job you run most.
Does Muse Glimmer run on Kali Linux?
Yes, through Ollama or llama.cpp, with the same memory needs. Inside a Kali virtual machine it is CPU-only and slow; run it on the host or a bare-metal install with an NVIDIA or AMD card and point tools in the VM at the host's Ollama port.
Is Muse Glimmer good for coding?
For agentic coding it is strong: 51.2 on SWE-Bench Pro and 76.0 on SWE-Bench Verified, close to Qwen3.6-27B. For terminal-heavy work Qwen leads. Reviewers note it does best with explicit, step-by-step instructions.
What is the Muse Glimmer context length?
131,072 tokens in the official files, and Unsloth's builds allow up to 262,144. Its sliding-window attention keeps the memory cost of long context low; a 130K context fit in 19.3 GB on an RTX 4090 in one test.
Will there be a smaller Muse Glimmer?
Meta has not announced one. As of late September 2026 the only release is the 30B, in full precision and 4-bit forms, plus the drafter and vision encoder. Any 8B or 3B version would be a guess.
If you did what Jake did, downloaded 18 GB on the promise of "runs on a single GPU" and got a model that types like it is thinking about every word, please don't take it personally. The claim was true for a card you do not have, and nothing on the page said which card. Glimmer is a genuinely good agent model on the right machine, and a fair one to skip on the wrong one. If anything here does not match what you see, tell me through the contact page. This guide will be updated when Meta ships a smaller Glimmer, when the Vulkan DFlash path settles, and as the Ollama tags change.
📌 If you keep one line from this page
Muse Glimmer is a dense 30B: every parameter has to be in fast memory, so 24 GB of VRAM or a 32 GB Mac is the real minimum.
Use the -dflash tag or the --spec-type draft-dflash flag, or you leave the advertised speed on the table.