Run JetBrains Mellum 2.1 Locally on Windows 11, Mac and Kali Linux: Free Install Guide (Ollama, IntelliJ, VS Code) and Why Qwen Still Wins

Logeshwaran
—
Run JetBrains Mellum 2.1 Locally on Windows 11, Mac and Kali Linux: Free Install Guide (Ollama, IntelliJ, VS Code) and Why Qwen Still Wins

Mellum 2.1 is JetBrains' new free coding model, released on October 8, 2026 under the Apache 2.0 license. It is a 12-billion-parameter model that only uses 2.5 billion parameters for each word it writes, which is why you can install it locally on a 16 GB Windows 11, Mac or Kali Linux laptop: the recommended download is an 8.1 GB file, and it runs fully offline. JetBrains taught it to work inside real code repositories, and its score on the SWE-bench Verified bug-fixing test jumped from 2.0 to 47.0 in four months. Here is the surprise, though, and it comes from JetBrains' own chart: Qwen3.5 9B still beats it on 12 of the 17 tests JetBrains published, including every hard agent benchmark. What Mellum 2.1 wins is coding puzzles and speed. Under heavy load it serves almost twice as many tokens as Qwen3.5 9B. And if you run it in Ollama on a normal laptop, it quietly gets a 4,000-token memory, far too small for coding work, unless you change one setting.

Jake runs a phone repair shop, and his friend Ethan, a developer, wrote him a little Python script years ago that books repairs and texts customers when their phone is ready. It mostly works. "Mostly" is the problem. When Ethan saw that JetBrains, the company behind IntelliJ and PyCharm, had released a free coding model that runs on a laptop, he had one question: could it look after Jake's script without sending the code, and the customers' phone numbers inside it, to anyone's cloud? Jake had a different question: "Is it as good as the one you pay for?" This page answers both, honestly. It covers what Mellum 2.1 is in plain English, what the benchmarks really say, what hardware you need, how to run it in Ollama, llama.cpp or LM Studio on Windows, Mac and Kali Linux, how to plug it into IntelliJ, PyCharm, VS Code and OpenCode, the mistakes that make it look broken, and where it fits on AWS.


Run JetBrains Mellum 2.1 Locally on Windows 11, Mac and Kali Linux: Free Install Guide (Ollama, IntelliJ, VS Code) and Why Qwen Still Wins

⚡ Quick Answer

• What is JetBrains Mellum 2.1? → A free, open coding model: 12B total parameters, 2.5B active, 131,072-token context, text only, Apache 2.0. It thinks before it answers. What it is.

• Is it better than Qwen? → Faster, and better at coding puzzles. Qwen3.5 9B still wins the hard agent tests on JetBrains' own chart. The honest numbers.

• Fastest way to run it → ollama run hf.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking-GGUF:Q4_K_M (8.1 GB). Raise Ollama's context to 64K first. Ollama steps.

• In IntelliJ or PyCharm → Settings, Tools, AI Assistant, Providers & API keys, then add Ollama or LM Studio. Use it for chat, not inline completion. IDE setup.

• Hardware → 16 GB of RAM runs the 8.1 GB file on the processor alone. A 12 GB graphics card holds it with room for context. Hardware tiers.

It is free to download and free to use commercially. The only thing it costs you is disk space and electricity.

New to running AI on your own machine? Our plain-English series on running AI locally explains models, parameters and file formats in about ten minutes. You can follow this page without it, though. Every step below starts from zero.

🧭 NEW HERE? READ THESE FIRST

New to local AI? These five pages make the rest of this one easy:

 Bookmark this; the context setting and the "looks broken" list are the parts you will come back for.

What JetBrains Mellum 2.1 is, in plain English

Mellum 2.1 is a large language model trained to write, read and fix code. You download it once, run it on your own computer or server, and it answers in your editor or terminal without sending anything to JetBrains or anyone else. Think of it as a junior developer who lives on your laptop: it can explain a function, find a bug, draft a fix and check its own work, as long as you give it the right tools.

The "12B, 2.5B active" part is the clever bit, and it is worth one minute. Mellum 2.1 is a mixture-of-experts model. Inside it are 64 small groups of parameters called experts. For every word it writes, a router picks the 8 experts best suited to that word and ignores the other 56. So the model holds 12 billion parameters, but each word only costs the work of about 2.5 billion. Jake understood it immediately: "So it's a shop with 64 technicians, and only 8 of them come to the bench for each job." Exactly. You still need room for all 64, which is why the download is 8.1 GB, but each job is quick.

The facts, from JetBrains' model card:

  • Size: 12 billion parameters in total, 2.5 billion active per token, 64 experts with 8 active.
  • Layers: 28, with grouped-query attention (32 query heads, 4 key-value heads).
  • Context: 131,072 tokens, roughly 100,000 words of code and conversation. Three out of every four layers only look back 1,024 tokens, which keeps the memory cost of long context low.
  • Thinking: it writes its reasoning inside <think>...</think> tags before the final answer. Only the "Thinking" version of 2.1 was released.
  • Input: text only. It cannot read screenshots or images.
  • License: Apache 2.0, free for personal and commercial use, with no sign-up gate on Hugging Face.
  • Hugging Face name: JetBrains/Mellum2.1-12B-A2.5B-Thinking, with official GGUF files in JetBrains/Mellum2.1-12B-A2.5B-Thinking-GGUF.

What changed from Mellum 2 is not the model's shape. The architecture is identical. Almost all of the work went into reinforcement learning, training by trial and reward. JetBrains built tasks in math, competitive programming, science, tool use and software engineering, and for the software tasks the model worked inside real repositories with a shell and file-editing tools, and was rewarded only when the tests passed. JetBrains says it ran millions of sandboxed attempts this way over the summer. The result, in their words, is a model that can explore a codebase, edit files and check its own changes, something Mellum 2 could not do well.

JetBrains names three jobs for it:

  • A worker inside coding agents. Finding the root cause of a failing test, drafting a fix and checking it, usually as one step in a bigger system.
  • A general reasoning assistant for everyday questions and step-by-step math.
  • Private, self-hosted deployment, so your code and data stay under your control.

That third job is the one that made Ethan sit up. Jake's script contains customer phone numbers. Running the model locally means none of that ever leaves the shop's laptop.

What Mellum 2.1 is not

Four things trip people up in the first hour, so here they are before you download anything:

  • It is not an autocomplete model. The gray "ghost text" that finishes your line as you type needs a special fill-in-the-middle model. Mellum 2.1 is a chat and agent model.
  • It is not a hosted service. There is no Mellum website to chat with and no Mellum subscription. You run it, or a server you control runs it.
  • It is not Junie. Junie is JetBrains' coding agent product. Mellum is a model. JetBrains has not said that Mellum 2.1 powers Junie or AI Assistant.
  • It cannot see. Paste a screenshot of an error and it gets nothing. Paste the error text instead.

Mellum, Mellum 2 and Mellum 2.1: the family tree

Search results mix up three generations of Mellum, so here is the whole family in one table. The dates come from JetBrains' Hugging Face pages and announcements.

ModelOpen releaseSizeWhat it is for
Mellum (original)April 2025 (used inside JetBrains products since 2024)4B, denseCode completion. The base model needs fine-tuning before use; Python, Kotlin and "all" fine-tunes followed.
Mellum 2Released in June 2026 (weights uploaded May 26)12B total, 2.5B activeFast general model for routing, RAG and sub-agents. Base, Instruct and Thinking versions.
Mellum 2.1October 8, 202612B total, 2.5B active (same shape)Coding agents and hard reasoning. Thinking version only.

The practical lesson: if a page or an Ollama tag says "mellum2" without the ".1", it is the June model, the one that scored 2.0 on SWE-bench Verified. If it says "mellum-4b", it is the 2025 completion model. Both are still on Hugging Face, and both are easy to download by mistake.

Mellum 2.1 benchmarks: the honest picture

JetBrains published 17 benchmark results, all measured by JetBrains with one pipeline, in thinking mode. That is useful, because every model was tested the same way, and it is also the main caveat: nobody outside JetBrains has reproduced them yet. Here is the whole table, with the higher score in each row being better, except HarmBench, where lower is better.

BenchmarkMellum 2.1Mellum 2Gemma 4 E4BQwen3.5 9B
Coding
LiveCodeBench v682.069.469.475.4
HumanEval+91.590.989.189.6
MBPP+79.475.470.969.8
Math
AIME 2025/202683.360.145.086.7
GSM-Plus88.387.187.491.4
Agentic
SWE-bench Verified47.02.023.050.0
Terminal-Bench 2.117.40.63.421.7
SWE-bench Pro28.00.04.038.0
Tool use
BFCL v462.349.652.558.5
WorkBench44.645.146.139.7
ToolHop49.146.739.952.0
Instructions and knowledge
IFEval90.679.590.892.4
GPQA Diamond64.651.053.177.8
MMLU-Redux87.886.084.989.5
MixEval-Hard46.441.641.450.3
Safety
XSTest (safe compliance)88.891.295.295.6
HarmBench (lower is better)8.521.538.16.6

Count the bold numbers and the story tells itself. Against its own predecessor, Mellum 2.1 wins 15 of 17 rows. Against Google's Gemma 4 E4B, it wins almost everything. Against Qwen3.5 9B, it wins 5: the three coding-puzzle tests and two of the three tool-use tests. Qwen3.5 9B takes the other 12, including all three agent benchmarks, the ones that most resemble "fix a real bug in a real repository."

So is JetBrains' release a disappointment? Not quite, and the reason is what each model costs to run. Qwen3.5 9B is a dense model: every one of its 9 billion parameters works on every word. Mellum 2.1 does the work of about 2.5 billion per word. On a laptop that difference is very noticeable, and JetBrains' speed chart shows it on a server too.

A few details matter if you compare these numbers with other charts. The agent tests ran in an open-source harness called Pi, version 0.73.1, with a shell and file tools, a 114,000-token context, up to 16,000 tokens per turn, and each model's default sampling settings (temperature 1.0 for Mellum 2.1). The non-agent tests used greedy decoding, which means the model always picked its single most likely next word. AIME is the average of the 2025 and 2026 exams, 30 questions each. JetBrains measured Mellum 2 again with this pipeline, so its numbers differ slightly from the Mellum 2 technical report.

How fast is Mellum 2.1?

JetBrains measured speed on a single NVIDIA H200, a data-center GPU, and published two claims. Under heavy load, Mellum 2.1 is the fastest model in the group and serves almost twice as many tokens as Qwen3.5 9B. For a single request, multi-token prediction makes it about 1.6 times faster.

There is a catch in that second number. Multi-token prediction needs an extra small "draft head" that guesses several words ahead, and JetBrains says that head for vLLM is still "coming soon." So the 1.6x speed-up is not something you can switch on today. The baseline speed is the same as Mellum 2's, which is already quick because of the 2.5B active parameters.

On a laptop, the useful rule of thumb is this: a mixture-of-experts model reads only its active experts for each word, so it generates text at roughly the speed of a small dense model, while needing the memory of a large one. That is why an 8.1 GB Mellum 2.1 file can feel quicker on a processor than a smaller dense 9B model. Your exact speed depends on your memory bandwidth more than your processor's core count.

Mellum 2.1 hardware: RAM, GPU and which file to download

JetBrains publishes five official GGUF files, the single-file format used by Ollama, llama.cpp and LM Studio. Each is one download. KLD and top-token match measure how close each smaller file stays to the full-precision model; lower KLD and a higher match are better.

FileSizeKLD vs BF16Top-token matchWho it is for
MXFP4_MOE7.0 GB0.11585.6%The smallest. Tight 16 GB laptops and 8 GB graphics cards.
Q4_K_M (recommended)8.1 GB0.07588.0%Most people. 16 GB laptops, 12 GB graphics cards.
Q6_K10.9 GB0.02093.9%16 GB graphics cards, 24 GB+ Macs.
Q8_012.9 GB0.00896.1%Effectively lossless. 24 GB cards, 32 GB machines.
BF1624.3 GBreferencereferenceResearchers. Almost nobody needs this at home.

The file size is not the whole memory bill. The model also needs working memory for the conversation, called the KV cache, and that grows with context length. Mellum 2.1 is unusually kind here: only one layer in four keeps the whole conversation, and the other three keep a sliding window of 1,024 tokens. A long coding session still costs memory, just much less than it would on a standard 12B model.

Here is what that means for real machines:

  • 8 GB of RAM, no graphics card: not realistic. Even the 7.0 GB file leaves nothing for Windows, your editor and the context.
  • 16 GB of RAM, no graphics card: the Q4_K_M or MXFP4_MOE file runs on the processor. Close the browser tabs you do not need. This is Jake's laptop, and it is enough for chat-style help.
  • 8 GB graphics card: the MXFP4_MOE file mostly fits, and the rest spills into system RAM. It works; it is just slower than a full fit.
  • 12 GB graphics card: Q4_K_M fits with room for a useful context. A comfortable home setup.
  • 16 GB graphics card: Q6_K with a generous context, or Q4_K_M with a very long one.
  • 24 GB card or 32 GB+ Mac: Q8_0, close to the full model, with room for agent-length conversations.

Apple Silicon Macs share one pool of memory between the processor and graphics, so the "graphics card" rows apply to your total memory minus what macOS and your apps are using. If you are shopping for a machine, our honest guide to laptops for local AI has the tiers.

How to install Mellum 2.1 locally with Ollama (Windows 11, Mac, Linux and Kali)

Ollama is the easiest way to install Mellum 2.1 locally, and the steps are the same on Windows 11, macOS and Kali Linux. There is no official "mellum2.1" entry in Ollama's own model library yet, but you do not need one: Ollama can pull GGUF files straight from Hugging Face, and JetBrains' GGUF page gives the exact command. If you are still choosing a local AI app, our comparison of Ollama, LM Studio, Jan and the rest explains the trade-offs.

Step 1: install or update Ollama

  • Windows: download OllamaSetup.exe from ollama.com, or run irm https://ollama.com/install.ps1 | iex in PowerShell.
  • Mac: download Ollama.dmg, or run curl -fsSL https://ollama.com/install.sh | sh in Terminal.
  • Linux and Kali: run curl -fsSL https://ollama.com/install.sh | sh. On Kali this also sets Ollama up as a background service.

Step 2: raise the context length before anything else

This is the step almost every guide skips, and it is the one that decides whether Mellum 2.1 looks clever or broken. Ollama chooses a default context length from the graphics memory it finds:

Your graphics memoryOllama's default contextEnough for coding work?
Under 24 GB (most laptops and gaming PCs)4,096 tokensNo
24 to 48 GB32,768 tokensFor chat, yes. For agents, tight.
48 GB or more256K tokensYes

Ollama's own documentation says coding tools and agents need at least 64,000 tokens. A thinking model makes the problem worse, because Mellum 2.1 writes its reasoning before its answer, and on a hard question the reasoning alone can run past 4,000 tokens. With the default, the model forgets the start of your question before it finishes thinking about it.

Fix it in one of two ways:

  • In the Ollama app: open Settings and move the context length slider to 64K.
  • On the command line: start the server with OLLAMA_CONTEXT_LENGTH=64000 ollama serve. This is the usual route on Linux and Kali; on Windows and Mac, the app's slider is simpler.

On a 16 GB machine without a graphics card, 64K is still workable with the Q4_K_M file thanks to the sliding-window layers, but if memory runs short, 32K is a sensible middle ground for chat-style use.

Step 3: pull and run Mellum 2.1

ollama run hf.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking-GGUF:Q4_K_M

The part after the colon picks the file. Swap Q4_K_M for MXFP4_MOE, Q6_K or Q8_0 to match the hardware table above. The first run downloads the file; later runs start in seconds. To download without starting a chat, use ollama pull with the same name.

Step 4: check what you actually got

While the model is loaded, run:

ollama ps

Two columns matter. CONTEXT should show the 64,000 you set, not 4,096. PROCESSOR shows how much of the model sits on your graphics card; "100% GPU" is ideal, a split like "40%/60% CPU/GPU" means part of it spilled into slower system memory. If the GPU column is empty when you expected it to be used, our guide to Ollama not using the GPU walks through the causes. If Ollama refuses to load the model at all, our guide to Ollama memory errors covers that.

Step 5: a first real question

Ethan's first test was the bug JetBrains uses in its own example, because it is small and has a definite answer:

Find the bug in this function and explain the fix:
def mean(xs): return sum(xs) / len(xs) - 1

A good answer spots two problems. The - 1 is subtracted from the finished average, because division happens before subtraction, so every result is one too small. And an empty list crashes with a division-by-zero error. You will see the model's reasoning appear first, inside the thinking tags, then the answer. That is normal for this model, and it is the reasoning that makes it good at bugs like this.

Ollama also exposes an OpenAI-compatible API at http://localhost:11434/v1, so any tool that speaks the OpenAI API can use Mellum 2.1 by pointing at that address and using the full hf.co/... name as the model.

Run Mellum 2.1 with llama.cpp or LM Studio

llama.cpp: one command, with more control

llama.cpp is the engine underneath Ollama and LM Studio, and running it directly gives you every knob. JetBrains' own command downloads the file and starts a server in one step:

llama-server -hf JetBrains/Mellum2.1-12B-A2.5B-Thinking-GGUF:Q4_K_M \
  --ctx-size 131072 --temp 0.6 --top-p 0.95 --top-k 20

On a laptop, --ctx-size 131072 asks for the full context and may not fit; --ctx-size 65536 is the realistic setting on 16 GB. The sampling values, temperature 0.6, top-p 0.95 and top-k 20, are the ones JetBrains uses in its examples.

Two newer llama-server options are made for thinking models like this one:

  • --reasoning-budget N caps how many tokens the model may spend thinking. The default, -1, is unlimited. A budget such as 4096 keeps simple questions quick, at some cost on hard ones.
  • --reasoning-format deepseek moves the thinking text into a separate reasoning_content field, so your app shows only the answer. The default, auto, usually does the right thing.

The server speaks the OpenAI API at http://localhost:8080/v1. Here is JetBrains' Python example, pointed at your own machine:

from openai import OpenAI

client = OpenAI(base_url="http://localhost:8080/v1", api_key="llama.cpp")

reply = client.chat.completions.create(
    model="JetBrains/Mellum2.1-12B-A2.5B-Thinking-GGUF",
    messages=[{"role": "user", "content": "Is 1024 a power of 2? Explain your reasoning."}],
    max_tokens=81920,
    temperature=0.6,
    top_p=0.95,
    extra_body={"top_k": 20},
)
print(reply.choices[0].message.content)

Notice max_tokens=81920. That number looks huge, and it is deliberate: the thinking counts toward the limit, and a small limit cuts the model off in the middle of its reasoning, before it ever reaches the answer. Tool calling works through the same server, because llama.cpp now uses the chat template stored inside the GGUF by default.

LM Studio: the no-terminal option

If you prefer clicking to typing, LM Studio works too:

  1. Open LM Studio, go to the search (Discover) tab and search for Mellum2.1.
  2. Pick JetBrains' GGUF repository and the Q4_K_M file, or the size that matches your hardware.
  3. When you load the model, raise the context length in the load settings, to 32K or 64K if your memory allows.
  4. Chat in the app, or start the local server from the Developer tab. It listens on port 1234 by default and speaks the OpenAI API.

If LM Studio refuses with an error when you load it, our guide to LM Studio "failed to load model" errors decodes each message. The usual cause with a model this size is simply memory: try the MXFP4_MOE file or a shorter context.

How to use Mellum 2.1 in IntelliJ IDEA, PyCharm and other JetBrains IDEs

Here is a slightly odd fact: JetBrains released Mellum 2.1, but there is no Mellum 2.1 switch in JetBrains IDEs. You connect it the same way you connect any local model, through AI Assistant's third-party providers. The steps, from the AI Assistant documentation for current IDE versions:

  1. Start the model first. Have Ollama or LM Studio running with Mellum 2.1 downloaded, using the steps above.
  2. In your IDE, open Settings | Tools | AI Assistant | Providers & API keys.
  3. In the Third-party AI providers section, choose Ollama or LM Studio. For llama.cpp or another server, choose the OpenAI-compatible option instead.
  4. Enter the address: http://localhost:11434 for Ollama, http://localhost:1234 for LM Studio, or http://localhost:8080/v1 for llama-server. Click Test Connection.
  5. Enable third-party AI providers and click Apply.
  6. Open AI Chat. Mellum 2.1 appears in the model selector under its provider's name.

On older IDE versions the menu may read "Models & API keys" instead of "Providers & API keys." It is the same page.

Which AI Assistant features to give it

Further down the same settings page, Models Assignment decides which model does which job. Local models are never assigned automatically, so you choose:

  • Core features: in-editor code generation, commit messages, the default chat model and similar. Mellum 2.1 is a good fit here.
  • Instant helpers: chat titles, name suggestions and context collection. These need speed more than brains; Mellum 2.1 works, but its thinking makes it slower than a small non-thinking model for these tiny jobs.
  • Context window: the default for local models is 64,000 tokens. Make sure your Ollama or LM Studio context is at least as large, or the IDE sends more than the model can hold.

Leave inline code completion alone. AI Assistant configures completion separately, and it needs a fill-in-the-middle model, a type trained to complete code from both sides of the cursor. Mellum 2.1 is a chat model, so assigning it there gives poor or no suggestions.

Two limits are worth knowing before you plan around it. AI Assistant does not currently call tools from your MCP servers when it uses a local model. And some features may be unavailable with third-party models unless a JetBrains AI subscription is active alongside them. On cost, JetBrains' free AI tier, introduced with the 2025.1 IDEs, includes local model support and unlimited code completion, so connecting Mellum 2.1 does not need a paid plan.

Use Mellum 2.1 in VS Code and OpenCode

VS Code with Copilot Chat

Yes, a JetBrains model works in Microsoft's editor. VS Code's chat can use models from other providers, and Ollama is one of the built-in ones:

  1. Start Ollama with Mellum 2.1 pulled and the context raised, as above.
  2. In the Chat view, open the model picker and choose Manage Models, or run Chat: Manage Language Models from the Command Palette.
  3. Add the Ollama provider and enable the Mellum 2.1 model.
  4. Pick it in the chat model selector.

Two things surprise people. Local models in VS Code work in chat and agents, but not in the gray inline completions. And a model only appears in Agent mode if VS Code sees that it supports tool calling; if Mellum 2.1 is missing from the agent picker but present in chat, that is the reason. Business and Enterprise Copilot plans may also need an administrator to allow models from other providers.

OpenCode, the free terminal agent

OpenCode is a free, open-source coding agent that reads your project, edits files and runs commands. It is where Mellum 2.1's agent training is most useful. Our OpenCode install guide covers the setup; to add Mellum 2.1, list it under an Ollama provider in opencode.json:

{
  "$schema": "https://opencode.ai/config.json",
  "provider": {
    "ollama": {
      "npm": "@ai-sdk/openai-compatible",
      "name": "Ollama (local)",
      "options": { "baseURL": "http://localhost:11434/v1" },
      "models": {
        "hf.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking-GGUF:Q4_K_M": { "name": "Mellum 2.1" }
      }
    }
  },
  "permission": { "edit": "ask", "bash": "ask" }
}

The last line matters as much as the model. OpenCode lets the AI edit files and run shell commands without asking unless you change it, and a 12B model will sometimes make a confident wrong move. With "ask," you approve each one. The 64K context rule applies here too: it is the fix behind most "OpenCode doesn't work with Ollama" reports.

Mellum 2.1 as a coding agent: a real example

Ethan pointed OpenCode, running Mellum 2.1, at Jake's booking script and asked one thing: "Customers say they sometimes get two 'your phone is ready' texts. Find out why." Here is roughly how a run like that goes, and why the reinforcement learning matters.

An old-style model would read the question and guess. Mellum 2.1 behaves the way it was trained to: it lists the files, opens the one that sends texts, searches for where that function is called, and finds that it is called from two places, once when a repair is marked done and again when a technician closes the ticket. It then proposes a fix, a "notified" flag checked before sending, and, because OpenCode is set to ask, waits for Ethan to approve each edit. If there are tests, it runs them to check its own change. That loop of explore, edit, verify is exactly what the SWE-bench Verified jump from 2.0 to 47.0 measures.

It is also where the limits show. On a large, unfamiliar repository, a 12B model loses the thread sooner than a frontier model. The practical habits that help:

  • Give it a narrow job. "Find why the texts are sent twice" works far better than "improve this project."
  • Point it at the right folder. Start the agent in the subfolder that matters, not the root of a monorepo.
  • Keep the tests close. A model that can run a test checks itself. A model that cannot is guessing.
  • Use Git. Commit before every agent session, so any change is one command away from undone.

Serving Mellum 2.1 for a team with vLLM

For a shared server with a proper graphics card, vLLM is the standard choice, and JetBrains gives the commands. Without tool calling:

vllm serve JetBrains/Mellum2.1-12B-A2.5B-Thinking \
  --max-model-len 131072 \
  --reasoning-parser qwen3

With tool calling, which agents need:

vllm serve JetBrains/Mellum2.1-12B-A2.5B-Thinking \
  --max-model-len 131072 \
  --reasoning-parser qwen3 \
  --enable-auto-tool-choice \
  --tool-call-parser hermes

The full-precision weights are 24.3 GB, so plan on a GPU with clearly more than 24 GB of memory once the context is added; a 48 GB card is comfortable. On Kali or any recent Debian-based system, install vLLM inside a virtual environment, or pip refuses with "externally-managed-environment." Our guide to that pip error explains the right way.

Mellum 2.1 vs Qwen, Gemma 4, Claude, Copilot and Junie

"Mellum vs" is most of what people search, and the comparisons mix up models, products and services. Here is each one in plain terms.

Compared withWhat it isPick Mellum 2.1 whenPick the other when
Qwen3.5 9BOpen dense model, all 9B parameters work on every wordYou want speed, coding puzzles or many requests at onceYou want the best agent and knowledge scores at this size
Gemma 4 E4BGoogle's small open model, also sees images and hears audioThe job is codeYou need images or audio, or the smallest download
Qwen3.6Newer Qwen family, including larger mixture-of-experts modelsYour machine has about 16 GBYou have more memory and want a bigger model; JetBrains did not test against it, so compare on your own code
Claude and Claude CodeAnthropic's hosted frontier models and its coding agentCode must stay on your machine, or the job is small and frequentThe task is large, unfamiliar or high-stakes
GitHub CopilotA paid service with hosted models and inline completionYou want free, private chat; you can even add Mellum to VS Code's chatYou want ghost-text completion and frontier models with no setup
JunieJetBrains' coding agent product, not a modelNot a real either-or: Mellum is a model you plug into agentsYou want JetBrains' managed agent inside the IDE

The honest summary Ethan gave Jake: "It's not as smart as the one I pay for. It's free, it's fast, and your customers' phone numbers never leave this laptop. For your script, that's the right trade." For a large commercial codebase with a deadline, the right trade may be the hosted model. For a private script, a sub-agent doing one narrow job thousands of times, or a company that cannot send code outside, Mellum 2.1 is now a serious option.

If you would rather pay for a frontier agent, our guide to Kimi Code plans and setup shows what the cloud end of the scale costs. If you want a bigger open model locally, our guide to running Qwen3.8 on Windows and Kali covers that family.

Why Mellum 2.1 looks broken: the common mistakes

Most "Mellum 2.1 is bad" first impressions come from setup, not the model. Here is the list, roughly in order of how often they bite.

The 4K context in Ollama

Covered above, and it deserves the top spot. On any machine with less than 24 GB of graphics memory, Ollama gives the model 4,096 tokens unless you change it. The model then forgets the start of your files, loops, or stops mid-thought. Set 64K, then confirm with ollama ps.

An output limit that cuts off the thinking

Many apps cap a reply at a few thousand tokens. A thinking model spends tokens on reasoning first, so the cap arrives before the answer and you get a half-finished thought with no conclusion. Raise the maximum output tokens in your app or API call; JetBrains' examples use 81,920. If you want shorter thinking instead, cap the thinking itself with llama-server's --reasoning-budget.

The wrong Mellum

Ollama's library already has Mellum models published under JetBrains' name, such as mellum2-thinking-q4_k_m and mellum2-instruct-q4_k_m. Those are the June Mellum 2 models, the ones that scored 2.0 on SWE-bench Verified. For 2.1, use the hf.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking-GGUF name until an official 2.1 entry appears in the library.

Sampling settings from another model

Copying a config from another model can leave Mellum 2.1 with a temperature that is too high or too low. JetBrains' examples use temperature 0.6, top-p 0.95 and top-k 20. Its agent benchmarks used the model's default sampling, temperature 1.0. Start with JetBrains' values and change one setting at a time.

Using it for inline completion

Mellum 2.1 is not a fill-in-the-middle model. Assigned to inline completion in a JetBrains IDE or a VS Code extension, it gives poor suggestions or none. Use it for chat, explanations, fixes and agents.

Tool calls that come back as plain text

In vLLM, tool calling needs both --enable-auto-tool-choice and --tool-call-parser hermes; without them, the tool call arrives as ordinary text and your agent never acts on it. In VS Code, a model without detected tool support is hidden from Agent mode. In JetBrains AI Assistant, MCP tools are not called with local models at all, which is a limit of the IDE, not the model.

Thinking text in your output

If the <think> block shows up where your app expected a clean answer, tell the server to separate it: --reasoning-parser qwen3 in vLLM, or --reasoning-format deepseek in llama-server.

Downloading the 24 GB file by accident

Clicking the first file in the GGUF list gets you BF16, 24.3 GB. Almost everyone wants Q4_K_M at 8.1 GB. Check the name before a slow download eats your evening.

"Out of memory" on a 16 GB machine

The model file, the context and everything else on the computer share the same 16 GB. Close the browser, use MXFP4_MOE or Q4_K_M, and try 32K context before 64K. If Ollama still refuses, the memory-error guide linked above lists every fix.

Mellum on AWS: Bedrock Marketplace, SageMaker JumpStart or your own GPU

Mellum has a longer history on AWS than most people expect, and the three generations are in three different places. New to these services? Our Bedrock vs SageMaker AI explainer sorts them out first.

  • Amazon Bedrock Marketplace (since September 2025): JetBrains listed its production code-completion Mellum there. The model itself is free; you pay only for the GPU instance you deploy it on. This is the original completion model, not 2.1.
  • Amazon SageMaker JumpStart (since August 10, 2026): AWS added Mellum2-12B-A2.5B-Thinking to the JumpStart catalog, deployable from the SageMaker console or the SageMaker Python SDK. That is Mellum 2, the June model.
  • Mellum 2.1 today: AWS's announcements so far cover Mellum 2, not 2.1. To run 2.1 on AWS this week, host it yourself on a GPU instance with vLLM, using the commands above.

For self-hosting, size the GPU from the file you serve. The 24.3 GB full-precision weights plus context need a card with clearly more than 24 GB of memory; a 48 GB card such as the NVIDIA L40S in Amazon EC2 G6e instances is comfortable for vLLM. If a 24 GB card is what you have, serve the 12.9 GB Q8_0 GGUF with llama-server instead. Whichever you choose, keep the server inside your VPC and put authentication in front of it: an open model endpoint on the internet is an open invitation. Our explainer on Hugging Face models on AWS covers deployment patterns and costs in more depth.

JetBrains Mellum 2.1: frequently asked questions

What is JetBrains Mellum 2.1?

Mellum 2.1 is a free, open coding model from JetBrains, released October 8, 2026 under Apache 2.0. It is a 12B mixture-of-experts model with 2.5B active parameters, a 131,072-token context and a thinking mode, trained with reinforcement learning to work inside real code repositories.

Is Mellum 2.1 free?

Yes. The weights are free to download and the Apache 2.0 license allows commercial use. Running it locally costs nothing beyond your hardware. In JetBrains IDEs, the free AI tier includes local model support, so connecting it needs no paid plan.

How do I run Mellum 2.1 in Ollama?

Raise Ollama's context length to 64K in the app settings or with OLLAMA_CONTEXT_LENGTH=64000, then run ollama run hf.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking-GGUF:Q4_K_M. Check ollama ps to confirm the context and GPU use.

How much RAM does Mellum 2.1 need?

The recommended Q4_K_M file is 8.1 GB, so 16 GB of RAM runs it on the processor alone. A 12 GB graphics card holds it with room for context. The smallest file, MXFP4_MOE, is 7.0 GB; the full BF16 model is 24.3 GB.

Is Mellum 2.1 better than Qwen3.5 9B?

Not overall. On JetBrains' own 17 benchmarks, Mellum 2.1 wins 5, mainly coding puzzles and tool calling, and Qwen3.5 9B wins 12, including all three agent tests. Mellum 2.1 is faster, serving almost twice as many tokens under load.

What is the difference between Mellum 2 and Mellum 2.1?

The architecture is the same, 12B total and 2.5B active. Mellum 2.1 was retrained mostly with reinforcement learning in real repositories, which lifted SWE-bench Verified from 2.0 to 47.0. Only a Thinking version of 2.1 was released; Mellum 2 also has Base and Instruct.

Can I use Mellum 2.1 in IntelliJ IDEA or PyCharm?

Yes. Run it in Ollama or LM Studio, then open Settings, Tools, AI Assistant, Providers and API keys, choose the provider under Third-party AI providers, test the connection and apply. It then appears in AI Chat and can be assigned to core features.

Can Mellum 2.1 do code completion?

Not well. Inline completion needs a fill-in-the-middle model, and Mellum 2.1 is a thinking chat and agent model. The original 4B Mellum was the completion model. Use 2.1 for chat, fixes and agents.

Can I use Mellum 2.1 in VS Code?

Yes, in chat. Run it in Ollama, then use Manage Models in VS Code's chat to add the Ollama provider. Local models work in chat and agents but not inline completions, and Agent mode needs tool-calling support.

Is Mellum 2.1 as good as Claude Code or Copilot?

No, and it is not meant to be. Hosted frontier models handle large, unfamiliar tasks better. Mellum 2.1's strengths are privacy, zero cost per use and speed, which suit private code, narrow repeated jobs and sub-agents.

Does Mellum 2.1 power Junie?

JetBrains has not said so. Junie is JetBrains' coding agent product, and Mellum is a model. JetBrains' announcement positions Mellum 2.1 as an open model for your own agents and self-hosted deployments.

Does Mellum 2.1 support tool calling?

Yes. JetBrains documents vLLM tool calling with --enable-auto-tool-choice and --tool-call-parser hermes. llama.cpp's server supports tool calls through the chat template in the GGUF. Its BFCL v4 tool-calling score is 62.3.

Can Mellum 2.1 read images or screenshots?

No. It is a text-only model. Paste error messages and code as text instead of screenshots.

How do I install Mellum 2.1 locally on Windows 11?

Install Ollama from ollama.com, open its Settings and set the context length to 64K, then run ollama run hf.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking-GGUF:Q4_K_M in a terminal. The 8.1 GB download happens once. 16 GB of RAM is enough without a graphics card.

Does Mellum 2.1 work offline?

Yes. After the one-time download, Ollama, llama.cpp and LM Studio run it entirely on your machine with no internet connection, which is the point of a local install for private code.

Where can I download Mellum 2.1?

On Hugging Face: JetBrains/Mellum2.1-12B-A2.5B-Thinking for the full weights, and JetBrains/Mellum2.1-12B-A2.5B-Thinking-GGUF for the Q4_K_M, MXFP4_MOE, Q6_K, Q8_0 and BF16 files used by Ollama, llama.cpp and LM Studio.

Is Mellum 2.1 on Amazon Bedrock?

Not yet. The original completion Mellum is on Bedrock Marketplace and Mellum 2 Thinking is in SageMaker JumpStart. To run 2.1 on AWS now, host it yourself with vLLM on a GPU instance with more than 24 GB of memory.

Why does Mellum 2.1 stop in the middle of thinking?

The output limit is too small. Its reasoning counts toward max tokens, so a low cap ends the reply before the answer. Raise the limit, JetBrains uses 81,920 in its examples, or cap the thinking with llama-server's --reasoning-budget.

What languages does Mellum 2.1 know?

JetBrains lists English as the language of the model card, and it was trained on code and natural language. It handles mainstream programming languages; test it on your own stack before relying on it for a niche one.

Can I run Mellum 2.1 on Kali Linux?

Yes. Install Ollama with curl -fsSL https://ollama.com/install.sh | sh, start it with OLLAMA_CONTEXT_LENGTH=64000 ollama serve, and run the hf.co command. For vLLM, install inside a Python virtual environment to avoid the externally-managed-environment error.

By the end of the evening, Jake's double-text bug was fixed, the change was committed, and the customers' phone numbers had never left the shop. Jake asked whether he should cancel anything. Ethan laughed: "Keep the paid one for the hard days. This one is for every other day." That is the honest place for Mellum 2.1. It is not the smartest coding model you can use. It may be the smartest one that is free, fast, private and small enough to live on the laptop you already own, and four months ago it could not fix a bug at all.

 If you keep one line from this page

Mellum 2.1 is fast and free, Qwen is still smarter, and Ollama's 4K default breaks both.

Download Q4_K_M, raise the context to 64K, and give it narrow jobs with tests close by.

Revision note. Written October 9, 2026, the day after JetBrains released Mellum 2.1. If you have a small script you have been meaning to fix for years, this is a friendly, free way to start.

Related