Aleph Alpha Kolibri: Specs, Benchmarks, and How to Run It Locally (Mac, Windows, Kali)

Logeshwaran
—

Aleph Alpha Kolibri (Kolibri-1) is an open-weight AI model released on October 3, 2026 by Aleph Alpha, the German AI company. It is a 78-billion-parameter mixture-of-experts model that uses only 3.46 billion parameters per token, reads up to 1 million tokens of context, speaks German and English, and ships under the Apache 2.0 license, so businesses can use it commercially for free. Its own benchmark runs put it ahead of Qwen3.6-35B-A3B on math and competitive coding, and behind it on tool use and software engineering. Can you run Kolibri locally? Yes, with caveats: community MLX builds run it on a Mac with 36 GB or more of memory, and a patched llama.cpp runs it on a PC, even on the CPU alone, with 64 GB of RAM. Ollama and LM Studio cannot load it yet. Every step, every size and every honest limit is below.

Jake read the headline over Ethan's shoulder: "Seventy-eight billion parameters. So that's a server-room job, right?" It is the reasonable assumption, and for most models that size it is true. But Kolibri is built so that each word it writes only wakes up 3.46 billion of those parameters, a small fraction of the whole. Ethan pulled up the first test reports from people who had tried it on their own machines. One had it writing at 12 to 15 words a second on an ordinary desktop processor, with no graphics card involved at all, because the computer had 128 GB of memory to hold the full model. Another had it running at more than 50 a second on an M1 Max MacBook Pro from 2021. "So the catch isn't speed," Ethan said. "The catch is memory. You have to fit all 78 billion in, even though only a sliver works at a time." This page explains what Kolibri is, how good it actually is, who should care, and exactly how to run it on a Mac, Windows or Kali Linux, or the official way on a rented GPU.

⚡ Quick Answer

• What it is → 78B open model, 3.46B active, 1M context, Apache 2.0. Specs.

• How good → strong at math and reasoning, weaker at agentic coding. Benchmarks.

• On a Mac → MLX 2-bit to 4-bit builds, 36 GB to 64 GB of memory. Steps.

• On a PC or Kali → patched llama.cpp, 47.5 GB file, 64 GB of RAM. Windows and Kali.

Ollama or LM Studio? Not yet. What to do instead.

🧭 NEW HERE? READ THESE FIRST

New to running AI models on your own computer? These five pages pair with this one:

📌 Bookmark this; Kolibri needs memory more than speed.

What is Aleph Alpha Kolibri? The specs in one table

Kolibri is German for hummingbird: a small bird with very fast wings, which is the idea behind a large model that only moves a small part of itself at a time. The official name on Hugging Face is Aleph-Alpha/Kolibri-1.

SpecificationKolibri-1
ReleasedOctober 3, 2026
MakerAleph Alpha (Heidelberg, Germany)
Total parameters78.1 billion
Active per token3.46 billion
ArchitectureMixture-of-experts: 50 layers, 384 experts per layer, 6 chosen per token, plus 1 shared expert
Context window1,048,576 tokens (262,144 recommended for efficient serving)
LanguagesGerman and English
ReasoningBuilt-in thinking mode with effort none, low, medium or high
Tool callingYes, in the standard OpenAI-style tools format
WeightsFP8 (about 78 GB); a full-precision BF16 version is also published
LicenseApache 2.0 (commercial use allowed)
Knowledge cutoffJune 18, 2026

The release is a single instruction-tuned model, ready to chat, reason and call tools. There is no smaller sibling yet, so the choice is about which quantization (compressed copy) of the same model you run, not which size.

Who is Aleph Alpha? From Luminous to Kolibri

If the name rings a bell, it is probably from 2023, when Aleph Alpha was Europe's best-funded hope in the race to build large language models. Founded in Heidelberg in 2019 by Jonas Andrulis and Samuel Weinbach, it built the Luminous family of models and raised a round of more than half a billion euros in 2023 from backers including Bosch, SAP and the Schwarz Group, the owner of Lidl and Kaufland.

Then the company changed course. In 2024 it stepped back from competing head-on with the largest American and Chinese labs, stopped marketing the Luminous API to new customers, and launched PhariaAI, a platform for governments and companies to run and govern AI on their own terms. That is why people still search "is Aleph Alpha dead": the model maker had seemingly stopped making models.

Kolibri is the answer to that question. It is a new model, trained from scratch, built specifically for what Aleph Alpha calls sovereign, mission-critical work: public administration, industry, aerospace and other regulated fields where organizations need to know exactly how a model was built and run it on their own hardware.

The other big change is Cohere. In April 2026, the Canadian AI company Cohere announced plans to combine with Aleph Alpha, and on September 16, 2026 the two signed a definitive business combination agreement. The combined company is planned to have dual headquarters in Toronto and Berlin, keep Heidelberg as a research center, and run sovereign AI on STACKIT, the Schwarz Group's cloud. The deal still needs regulatory approval and is expected to close later in 2026. Kolibri carries the Aleph Alpha name and was released under it, two and a half weeks after the signing.

Is Aleph Alpha Kolibri open source? The license, plainly

Kolibri is open-weight under Apache 2.0, one of the most permissive licenses there is. In practice you can:

  • download the weights free of charge,
  • run them on your own hardware or in your own cloud,
  • use them in commercial products and services,
  • fine-tune them and publish your own versions,

as long as you keep the license and notices with anything you redistribute. That is a real difference from models released under "community" or research licenses with user caps or non-commercial clauses, such as the OpenJev model we covered this week.

Two fine-print points matter for anyone deploying it in a business. First, the Apache 2.0 grant covers the weights and configuration files; Aleph Alpha says it does not extend to its underlying code, training methods or other intellectual property. Second, Aleph Alpha publishes a list of prohibited uses (illegal activity, weapons, malicious code and uses banned under Article 5 of the EU AI Act, among others). Whether that list binds you depends on your lawyer's reading, but if your use is on it, Kolibri is the wrong tool anyway. Aleph Alpha is a signatory of the EU's General-Purpose AI Code of Practice and has published a technical report and a training-data summary, which is more disclosure than most open-weight releases offer.

So, is it open source in the strict sense? The weights are open; the training code and full dataset are not. "Open-weight with an Apache 2.0 license" is the accurate description, and for almost every practical purpose, running and building on it, that is what you need.

78 billion parameters, 3.46 billion at work: how that is possible

Most chatbots you have used are dense models: every parameter takes part in producing every word. Kolibri is a mixture-of-experts (MoE) model. Inside each of its 50 layers sit 384 small specialist networks, the "experts". For every token it reads or writes, a small router picks the six experts best suited to that token, plus one shared expert that always runs. The other 378 sit idle for that token.

Think of it as a very large hospital. All the specialists are on the payroll and in the building, which is why the building has to be big. But any one patient only sees the handful of doctors they need, which is why the waiting time is short. For a model, the "building" is memory and the "waiting time" is compute. That gives Kolibri two properties that surprise people:

  • It is fast for its size. Doing the work of a 3.46B model per token means it can write quickly even on hardware that would crawl with a dense 70B model. That is why a desktop CPU can run it at a usable speed.
  • It still needs all of its memory. Every expert must be loaded, because the next token might need any of them. A compressed 4-bit copy still takes about 44 to 48 GB.

Kolibri also mixes two kinds of attention, in a 4-to-1 ratio: most layers only look at a sliding window of recent text, and every fifth layer looks at everything. That keeps long documents affordable, and it is a big part of how the model handles a context window of up to a million tokens.

Kolibri benchmarks: what the numbers say

All the scores below are Aleph Alpha's own, run on its own test harnesses. No independent lab has reproduced them yet, so treat them as a strong first signal, not a final verdict. With that said, here is how Kolibri compares with two of the most relevant open models: Qwen3.6-35B-A3B, which also activates about 3 billion parameters per token, and NVIDIA's much larger Nemotron 3 Super (120B total, 12B active).

Benchmark (what it tests)KolibriQwen3.6-35B-A3BNemotron 3 Super
AIME 2026 (competition math)96.091.090.4
GPQA Diamond (graduate science)84.383.478.0
LiveCodeBench v6 (competitive coding)85.982.582.0
SWE-Bench Verified (fixing real code)66.473.860.2
BFCL v4 (function calling)61.467.261.0
Tau2-Bench Telecom (agent tasks)94.799.168.1

The pattern is clear. Kolibri is a reasoner: on math, science questions and competition-style coding it beats a similar-sized rival and a model with more than three times its active parameters. On agentic work, where the model must call tools reliably and fix real codebases step by step, Qwen3.6 is ahead, sometimes clearly. Aleph Alpha also reports Kolibri at 96.9 on AIME 2025, against 91.7 for Nemotron 3 Super and 79.8 for Mistral Small 4, and 92.7 on HumanEval+ against Nemotron's 94.7.

Overall, Aleph Alpha gives Kolibri an average of 75.5 across its English test suite and 70.8 across its German one. The German suite matters, because very few model makers test in German at all, and those that do usually run translated English tests.

The hallucination number worth noticing

One result stands out for business use. On a grounding test that measures how often a model avoids confidently wrong answers (AA-Omniscience's non-hallucination rate), Kolibri scores 44.0%, against 13.9% for Nemotron 3 Super and 56.7% for Qwen3.6. Aleph Alpha trained it with a technique it calls the Merlin-Arthur protocol to make it more willing to say "I don't know". It is not the best score in the table, but for a model meant for public administration, being trained to abstain rather than invent is the right instinct. Your own testing on your own documents is still the only number that counts.

How Kolibri was trained: the data and the compute

Aleph Alpha published more about Kolibri's training than most labs share, and the numbers explain its strengths. Training ran in three stages: a broad pre-training on 20 trillion tokens, a mid-training stage of 3.44 trillion tokens heavy on reasoning and tool use, and a short long-context stage of 201 billion tokens to stretch it to long documents. The mix of material changed sharply between stages:

Kind of dataPre-trainingMid-trainingLong-context
English web and documents43.4%1.3%15.8%
German web and documents23.4%2.1%2.6%
English instructions and reasoning11.9%31.4%11.3%
Code13.6%16.3%13.2%
Agentic code and tool useabout 0%15.4%5.6%
Science and math questions1.8%25.4%8.9%
Scanned PDFs (text recognized)0%0%33.3%

Two things jump out. A quarter of the mid-training stage was science and math questions, which matches Kolibri's strong math and science scores. And a third of the long-context stage was recognized text from scanned PDFs: long, messy, real documents of the kind public offices and companies actually hold. Tool use only arrived in mid-training, which may be part of why its agentic scores trail rivals trained on more of it.

The compute was substantial for a European lab: 768 NVIDIA B200 GPUs, 21 days of pre-training (about 392,000 GPU-hours), five more days of mid-training and 13 hours of long-context training, about 6.4 times ten to the 23rd power operations in total. Fine-tuning then added 4,000 steps of supervised training and 1,000 steps of reinforcement learning.

Where Kolibri falls short

Being honest about the weak spots saves you a wasted weekend:

  • Agentic coding. If you want a model to drive an editor, run tests and fix a repository on its own, the SWE-Bench and tool-calling numbers say Qwen3.6 is the better small-active-parameter choice today.
  • Only two languages. Kolibri is trained for German and English. It will often understand French, Spanish or Hindi, but it is not tuned for them, and multilingual models will do better.
  • No image input. Kolibri reads and writes text only. For documents with charts or scans, you need a separate step to turn them into text first.
  • Heavy memory needs. The low active-parameter count makes it fast, not small. A 16 GB laptop cannot run it at any quantization.
  • Young tooling. On release day, Ollama, LM Studio, stock llama.cpp and stock mlx-lm could not load it. Everything local today depends on community builds, covered below.

Why a German-first model matters

About a quarter of Kolibri's training text was German, far more than in most open models, where German is a small slice of the web data. Aleph Alpha also built its tokenizer (the part that chops text into pieces) to handle German efficiently: it packs about 4.7 bytes of German text into each token, compared with about 4.2 bytes of English. In plain terms, long German compound words cost fewer tokens than they do in many English-first models, which means faster answers and more German text fitting in the context window.

For a German municipality, insurer or engineering firm, this is the point. Legal and technical German is dense, formal and full of compound nouns, and English-first models often answer it in slightly stilted German. Kolibri was also trained on German reasoning specifically, with rewards for staying consistent when it thinks in German. If your documents, staff and customers work in German, it deserves a place on your shortlist even where its English scores trail a rival.

The 1-million-token context, honestly

Kolibri accepts up to 1,048,576 tokens, roughly 700,000 to 800,000 English words or several thick books at once. Three details keep that number in perspective:

  1. It was trained up to 256,000 tokens. The model was adapted to long documents at 256K, and the serving settings extend it to 1M. Aleph Alpha itself recommends serving at 262,144 for efficiency.
  2. Quality drops with length. On the RULER long-context test, Kolibri scores 63.2 at a full million tokens. At 256K, on the HELMET test, it scores 85.4. Both are respectable, but the model is far more dependable in the first quarter-million tokens.
  3. Memory grows with length. The weights are only part of the bill; the working memory for a long conversation (the KV cache) grows with every token. On a laptop, plan on contexts of 8,000 to 32,000 tokens, not a million.

For most real jobs, such as a long contract, a year of meeting notes or a large codebase folder, 256K is already more than enough.

Reasoning effort: none, low, medium, high

Kolibri can think before it answers, and you choose how much. Each request can set a reasoning effort of none, low, medium or high. Higher effort means a longer hidden chain of thought: better answers on math, logic and tricky questions, but slower replies and more tokens used. A sensible default:

  • none for chat, rewriting, translation and simple lookups,
  • low or medium for everyday questions, summaries and drafting,
  • high for math, planning, debugging and anything where a wrong answer is costly.

On a slow local machine, the effort setting is the biggest speed lever you have. A high-effort answer on a CPU can take minutes because the model writes a long reasoning trace first.

Can you run Kolibri locally? Every route, compared

Here is every way to run Kolibri as of October 4, 2026, from the official route to the community builds that appeared within a day of release.

RouteDownloadHardwareStatus
Official FP8 with vLLMAbout 78 GB1× H200, B200 or B300, or 2× A100 80 GB or H100Supported, with Aleph Alpha's plugin
MLX 4-bit (Mac)44.3 GBApple Silicon, 64 GB or moreCommunity build, custom launcher
MLX 3-bit (Mac)About 35 GBApple Silicon, 48 GB or moreCommunity build
MLX 2-bit (Mac)25.5 GBApple Silicon, 36 GB or moreCommunity build, lowest quality
GGUF Q4_K_M (llama.cpp)47.5 GB64 GB RAM or more; CPU testedNeeds a patched llama.cpp
NVFP4 W4A16 (vLLM)About 47 GBRecent NVIDIA GPUsExperimental; not validated
Ollama, LM Studio, Jann/an/aNot yet; they need the new architecture added

The reason for the patches is the same one we hit with other brand-new designs: runners only understand the model architectures they have code for, and Kolibri's ("kolibri1") is new. Community developers ported Aleph Alpha's open vLLM plugin to MLX and llama.cpp within a day. Those ports are impressive, but they are one person's code, tested on a few machines, so expect rough edges and check the repository pages for updates before you start.

How much RAM and VRAM does Kolibri need?

Memory decides everything. Use this as your rule of thumb, with headroom for the operating system and a working context:

Your machineCan it run Kolibri?Best option
Laptop or PC with 8 to 32 GB RAMNoA smaller local model, or Kolibri on a rented GPU
Mac with 36 GB unified memoryJustMLX 2-bit, short contexts
Mac with 48 GBYesMLX 3-bit
Mac with 64 GB or moreYes, wellMLX 4-bit (6-bit builds need more still)
PC with 64 GB RAM, any or no GPUYes, slowly to moderatelyPatched llama.cpp, Q4_K_M, on the CPU
PC with 128 GB RAMYes, comfortablyPatched llama.cpp, longer contexts
Single 24 or 32 GB gaming GPUNot on the GPU aloneUse system RAM with llama.cpp; GPU offload is untested
Workstation or cloud with 80 GB+ GPUsYes, the official wayvLLM with the FP8 weights

Measured speeds so far: the MLX builds generate about 52 to 56 tokens per second on an M1 Max, and the GGUF generates about 12 to 15 tokens per second on an AMD Ryzen 7 7800X3D CPU with 128 GB of RAM, using 8 threads. For comparison, most people read at about 4 to 5 words per second, so even the CPU speed is faster than you can read. Prompt reading on the CPU was 39 to 76 tokens per second, which means pasting a long document takes a while before the answer starts. Our honest guide to RAM and GPU tiers for local AI explains why memory, not processor speed, is usually the wall.

Run Kolibri on a Mac with MLX

Apple Silicon Macs are the easiest home for Kolibri today, because their unified memory lets the whole model sit in one pool the GPU can use. These steps use the community MLX build by velaia, which includes its own launcher because stock mlx-lm does not know the architecture yet.

  1. Check your memory: Apple menu > About This Mac. Pick the 4-bit build for 64 GB or more, 3-bit for 48 GB, 2-bit for 36 GB.
  2. Open Terminal and make a Python environment: python3 -m venv ~/kolibri && source ~/kolibri/bin/activate
  3. Install the tools: pip install -U mlx-lm huggingface_hub (version 0.32.0 or newer of mlx-lm is needed).
  4. Download the model (44.3 GB for 4-bit): hf download velaia/Kolibri-1-MLX-4bit --local-dir Kolibri-1-MLX-4bit
  5. Start a chat: python Kolibri-1-MLX-4bit/run.py chat --model Kolibri-1-MLX-4bit --temp 1.0 --top-p 0.97 --max-tokens 2048

For a single question instead of a chat, use run.py generate with --prompt "your question". Swap the repository name for velaia/Kolibri-1-MLX-3bit or velaia/Kolibri-1-MLX-2bit on smaller Macs. Lower bit widths lose some quality: the publisher measured perplexity (a measure of how surprised the model is by real text, lower is better) of 12.77 for German and 16.47 for English at 4-bit, rising to 14.01 and 17.46 at 2-bit. You will notice that as slightly less precise wording, more than as wrong facts, but it is real.

Other MLX conversions exist too, including 6-bit and mixed 4- and 8-bit builds for Macs with plenty of memory. Whichever you choose, close other memory-hungry apps first, and keep an eye on Activity Monitor's Memory Pressure graph: if it turns red, drop to a smaller build.

Run Kolibri on Windows with a patched llama.cpp

On Windows, the route is a community GGUF file plus a version of llama.cpp with Kolibri support patched in. The tested setup is CPU-only, which suits Kolibri better than it would a dense model. You need 64 GB of RAM or more and about 50 GB of free disk space.

The simplest path on Windows is WSL (Windows Subsystem for Linux), because the build steps are the same as on Linux. Open PowerShell as administrator, run wsl --install, restart, and open Ubuntu from the Start menu. WSL can use most of your RAM; if it is limited, raise the memory setting in the WSL Settings app. Then follow the Kali steps below, which work the same in Ubuntu.

If you prefer a native Windows build:

  1. Install Git for Windows, CMake, and Visual Studio Build Tools with the "Desktop development with C++" workload.
  2. Open the Developer PowerShell for VS and clone llama.cpp: git clone https://github.com/ggml-org/llama.cpp, then cd llama.cpp.
  3. Check out the commit the patch was made for: git checkout 836d571
  4. Download the patch file from the Hob-forge/Kolibri-1-GGUF repository on Hugging Face and apply it: git am C:\path\to\kolibri1-llama.cpp.patch
  5. Build: cmake -B build -DCMAKE_BUILD_TYPE=Release then cmake --build build --config Release -j --target llama-server llama-cli
  6. Download Kolibri-1-Q4_K_M.gguf (47.5 GB) from the same repository.
  7. Run the server: .\build\bin\Release\llama-server.exe -m Kolibri-1-Q4_K_M.gguf -c 32768 --jinja --temp 1.0 --top-p 0.97 --top-k 128
  8. Open http://127.0.0.1:8080 in your browser to chat.

The checkout matters: the patch was written against one specific llama.cpp commit and may not apply to newer code. If git am fails, you are on the wrong commit. The publisher tested only the CPU backend and only up to 8,192 tokens of context, so treat GPU offload and long contexts as experiments. If you are new to llama.cpp, our step-by-step llama.cpp guide from earlier this week covers the server, the web page and the common errors.

Run Kolibri on Kali Linux

Kali (and Ubuntu, Debian or WSL) builds llama.cpp in a few minutes. Check that you have 64 GB of RAM with free -h before you download anything.

sudo apt update
sudo apt install -y build-essential cmake git curl python3-pip
pip install -U huggingface_hub
git clone https://github.com/ggml-org/llama.cpp && cd llama.cpp
git checkout 836d571
hf download Hob-forge/Kolibri-1-GGUF --local-dir ~/models/kolibri
git am ~/models/kolibri/kolibri1-llama.cpp.patch
cmake -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build -j --target llama-server llama-cli
./build/bin/llama-server -m ~/models/kolibri/Kolibri-1-Q4_K_M.gguf \
  -c 32768 --jinja --temp 1.0 --top-p 0.97 --top-k 128

The hf download line pulls the whole repository, including the patch; if the patch has a different file name in your download, use that name in the git am line. If pip install complains about an externally managed environment on newer Kali releases, use pipx install huggingface_hub or a virtual environment instead. If apt cannot find a package, our guide to Kali's "Unable to locate package" error fixes the sources file in a minute.

Kali users often want a private model for security work: reading logs, explaining exploit code, drafting reports. Kolibri's long context suits pasting large log files, and nothing leaves the machine. Keep the server on 127.0.0.1 (the default) rather than exposing it on your network.

Set the reasoning effort in llama.cpp and MLX

The local servers speak the same OpenAI-style API as the official one. With llama-server running, you can set the effort per request:

curl http://127.0.0.1:8080/v1/chat/completions -H "Content-Type: application/json" -d '{
  "messages": [{"role": "user", "content": "Summarize this contract clause in plain German."}],
  "chat_template_kwargs": {"reasoning_effort": "low"}
}'

On a CPU, start with none or low. High effort can multiply the response time several times over, because every reasoning token has to be generated before the answer appears. Two more habits help on slow hardware:

  • Allow enough output tokens. Reasoning tokens count toward the reply limit. If answers stop mid-sentence at high effort, raise max_tokens in the request (or --max-tokens in the MLX launcher) rather than lowering quality.
  • Match effort to the question, not the session. Because effort is set per request, a chat app can use none for quick rewrites and high only when you ask it to check a calculation, which keeps the machine responsive most of the time.

Run Kolibri the official way: vLLM on a GPU server

For production or serious testing, Aleph Alpha's supported route is vLLM with its own plugin, on data-center GPUs. If you do not own one, rent a cloud instance with a single H200 or B200, or two 80 GB A100 or H100 cards, by the hour, and shut it down when you are done.

pip install 'aleph-alpha-inference>=1'
vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \
  --reasoning-parser kolibri1 \
  --tool-call-parser kolibri1 \
  --enable-auto-tool-choice

That serves the default 262,144-token context. For the full million, add --max-model-len 1048576 --hf-overrides '{"max_position_embeddings": 1048576}', and expect to need the larger GPU options. The recommended sampling settings are temperature 1.0, top-p 0.97 and top-k 128.

The server is OpenAI-compatible, so any app or library that talks to the OpenAI API can use it by changing the base URL. In Python:

from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
r = client.chat.completions.create(
    model="Aleph-Alpha/Kolibri-1",
    messages=[{"role": "user", "content": "ErklΓ€re kurz, was ein Mixture-of-Experts-Modell ist."}],
    extra_body={"chat_template_kwargs": {"reasoning_effort": "high", "enable_thinking": True}},
)
print(r.choices[0].message.content)

Tool calling works with the standard tools parameter: describe your functions in JSON schema, and the reasoning and tool-call parsers turn Kolibri's output into structured calls. Watch your bill on rented hardware: an idle GPU instance costs exactly as much per hour as a busy one, so stop it when you finish.

Add a chat window and document Q&A

The llama.cpp server includes a basic chat page, but most people want saved conversations and the ability to drop in documents. Open WebUI, a popular free interface for local models, can connect to any OpenAI-compatible server, which includes both llama-server and vLLM running Kolibri.

  1. Start Kolibri with llama-server (port 8080) or vLLM (port 8000) as shown above.
  2. Install Open WebUI, for example with Docker or pip install open-webui followed by open-webui serve.
  3. In Open WebUI, open Admin Panel > Settings > Connections, add an OpenAI API connection with the URL http://127.0.0.1:8080/v1 (or port 8000 for vLLM) and any placeholder key.
  4. Pick the Kolibri model in the chat window and start asking.

Open WebUI's document features let you upload files and ask questions about them, which plays to Kolibri's strengths in long German and English documents. Keep the context setting in llama-server (-c) large enough for the documents you load, and remember that a bigger context needs more memory.

What to actually use Kolibri for

The model is built for work, not for chit-chat, and the jobs it suits best follow from its training:

  • German paperwork. Summarizing official letters, drafting replies in formal German, explaining a tax notice or tenancy contract in plain language. This is the job its German training and tokenizer were built for, so it is the first thing to test it on.
  • Long-document questions. Asking questions across a long report, a set of meeting minutes or a policy manual, with the whole text in context instead of chopped into pieces.
  • Math and analytical reasoning. Checking calculations, working through statistics homework step by step, or planning with high reasoning effort.
  • Private analysis on Kali. Pasting large log files, explaining unfamiliar scripts, or drafting penetration-test reports, all without sending client data to an outside service.
  • A self-hosted company assistant. For organizations with GPU servers, a question-answering tool over internal documents, with human review, is the use Aleph Alpha designed it for.

Where it is the wrong pick: autonomous coding agents, image questions, and anything that needs languages beyond German and English. For those, the comparison table below points to better options.

Aleph Alpha Kolibri pricing: is it free?

The model itself is free: the weights cost nothing to download and the Apache 2.0 license charges nothing for commercial use. What costs money is running it:

  • On your own Mac or PC: free beyond electricity, if you already have the memory.
  • On rented GPUs: you pay the cloud provider by the hour for the instance, whatever you run on it.
  • Hosted by Aleph Alpha: no public price list or pay-as-you-go API was announced with the release. Enterprise deployments, including Aleph Alpha's PhariaAI platform and support, go through its sales team.

There is no free Kolibri chat website from Aleph Alpha at launch. Pages offering "Kolibri AI chat free" are third parties; check who runs them before pasting anything private.

Kolibri vs Qwen3.6 vs other open models: which should you run?

If you needPickWhy
German documents, reasoning, a model you can auditKolibriGerman-first training, published data summary, strong math and science
An agent that edits code and calls toolsQwen3.6-35B-A3BHigher SWE-Bench, BFCL and Tau2 scores, easier to run today
Something that runs in Ollama right now on 16 to 32 GBGemma 4 or a 9B modelKolibri does not fit; these do, with one command
Many languages, including Asian languagesA multilingual modelKolibri is tuned for two languages only
Images, scans or screenshots as inputA vision modelKolibri is text-only
EU data residency and a European supplierKolibriBuilt and trained in Germany and Finland, Apache 2.0, self-hostable

For most hobbyists with a normal laptop, the honest answer is that Kolibri is interesting to watch rather than something to install today. Our guide series to running AI locally lists the models that fit each hardware tier. For anyone with a 64 GB Mac or a workstation, and especially anyone working in German, it is worth an afternoon.

What "sovereign AI" means, and who should care

"Sovereign" is the word Aleph Alpha leads with, and it is more than marketing for one group of users. It means three concrete things here:

  • Where it was made. Kolibri was built in Germany and trained on infrastructure in Germany and Finland, under European law, on 768 NVIDIA B200 GPUs: 20 trillion tokens of pre-training over 21 days, then further training for reasoning and long documents.
  • What went in. Aleph Alpha published a training-data summary and a technical report describing its filtering: removing pirated sources, redacting personal data, de-duplicating, and checking for test contamination. Few open models disclose this much.
  • Who controls it after. With Apache 2.0 weights, a customer can run Kolibri entirely on its own servers, with no calls to any outside company, and keep it working even if the supplier changes direction or ownership.

For a ministry, a hospital group, a defense supplier or a bank in the EU, that combination answers questions their compliance teams ask about every AI tool. For an individual, it mostly means you know more than usual about where your model came from. Aleph Alpha also estimates the energy used for training at about 950 megawatt-hours, a disclosure the EU's code of practice encourages.

Safety, bias and limitations

Aleph Alpha is unusually direct about the risks, and they apply to every model of this kind:

  • Bias. The model reflects patterns in its training data, including cultural and political bias, which the company tried to reduce by filtering and by training on topics such as human dignity, democracy and the rule of law.
  • Outdated knowledge. It knows nothing after June 18, 2026. For current facts, give it the documents or connect it to search.
  • Hallucination. It is trained to abstain more often than many models, but it still makes things up. Verify anything important.
  • Harmful output. Built-in safeguards are not a guarantee; production systems need their own filtering and human review, especially for decisions about people.

The intended uses are telling: assistants, document drafting, question answering over an organization's own material, and decision support where a person makes the final call. Kolibri is meant to help people, not replace their judgment, and that is the right way to use any model.

Kolibri on Ollama and LM Studio: when?

As of October 4, 2026, there is no ollama pull kolibri, and LM Studio cannot load the GGUF. Both use llama.cpp (or MLX on Macs) under the hood, so they will be able to run Kolibri once support for its architecture lands in mainline llama.cpp and mlx-lm. The community patch and launcher are the first step; an upstream pull request and a release usually follow within weeks for popular models, though never on a guaranteed schedule.

Until then, you will see GGUF files appear on Hugging Face with Kolibri in the name. If LM Studio shows "unknown model architecture: kolibri1" when you load one, that is not a broken download; the runner simply lacks the code. We saw exactly the same pattern with other new architectures, explained in our guide to why LM Studio won't load some new models. We will update this page the day Kolibri reaches Ollama.

Aleph Alpha Kolibri: frequently asked questions

What is Aleph Alpha Kolibri?

An open-weight German and English AI model released on October 3, 2026: 78 billion parameters, 3.46 billion active per token, a 1-million-token context, under Apache 2.0.

Is Aleph Alpha Kolibri open source?

It is open-weight under Apache 2.0, so you can download, use commercially, modify and redistribute the weights. The training code and full dataset are not released.

Can I run Kolibri locally?

Yes, with enough memory. Community MLX builds run on Macs with 36 GB or more, and a patched llama.cpp runs it on PCs with 64 GB of RAM, even on the CPU alone.

How much RAM does Kolibri need?

About 25 GB for the smallest Mac build, 44 to 48 GB for 4-bit builds, and about 78 GB for the official FP8 weights, plus room for context and the system.

Is Kolibri on Ollama?

Not yet. Ollama and LM Studio need support for its new architecture first. Use the MLX builds or the patched llama.cpp in the meantime.

Can Kolibri run on a CPU without a GPU?

Yes. Because only 3.46 billion parameters work per token, a desktop CPU with 128 GB of RAM generated about 12 to 15 tokens per second with the Q4_K_M GGUF.

How fast is Kolibri on a Mac?

The community MLX builds generate about 52 to 56 tokens per second on an M1 Max, at 2-bit, 3-bit and 4-bit alike.

Is Kolibri better than Qwen3.6?

On Aleph Alpha's own tests, it beats Qwen3.6-35B-A3B at math, science and competitive coding, but trails on tool calling, agent tasks and SWE-Bench.

What languages does Kolibri support?

German and English. It is trained and evaluated in both, with an efficient German tokenizer.

Is Aleph Alpha Kolibri free?

The weights are free under Apache 2.0. Running it costs your own hardware or rented GPUs. No public hosted price was announced at launch.

What is Aleph Alpha pricing?

Aleph Alpha sells enterprise deployments and its PhariaAI platform through sales, without a public price list. Kolibri's weights themselves are free.

Is Aleph Alpha dead?

No. It stepped back from the frontier race in 2024 to focus on its PhariaAI platform, released Kolibri in October 2026, and signed an agreement to combine with Cohere.

What is the Aleph Alpha and Cohere deal?

A business combination agreement signed on September 16, 2026, creating a company with headquarters in Toronto and Berlin. It still needs regulatory approval and is expected to close later in 2026.

What happened to Aleph Alpha Luminous?

Luminous was Aleph Alpha's earlier model family. The company stopped marketing the Luminous API to new customers after its 2024 change of direction. Kolibri is its new model.

Does Kolibri really handle 1 million tokens?

It accepts 1,048,576 tokens, but it was trained to 256K and is more reliable there. It scores 63.2 on RULER at 1M and 85.4 on HELMET at 256K.

Can Kolibri read images?

No. Kolibri is text-only. Convert scans and charts to text first, or use a vision model.

What GPU do I need for the official Kolibri model?

At least one H200, B200 or B300, or two 80 GB A100 or H100 cards, for the FP8 weights with vLLM.

How do I set Kolibri's reasoning effort?

Pass reasoning_effort as none, low, medium or high in chat_template_kwargs with each request. Lower effort is much faster on local hardware.

Can I use Kolibri commercially?

Yes. Apache 2.0 allows commercial use, as long as you keep the license and notices with anything you redistribute.

Does Kolibri support tool calling?

Yes. It accepts the standard OpenAI-style tools parameter, and vLLM parses its tool calls with the kolibri1 parser.

Can I run Kolibri on Kali Linux?

Yes, with 64 GB of RAM: build llama.cpp at commit 836d571 with the community patch and run the Q4_K_M GGUF.

Why does LM Studio say unknown model architecture kolibri1?

LM Studio does not have code for Kolibri's new architecture yet. The download is fine; wait for an update or use the patched llama.cpp.

Is Kolibri good at coding?

Good at competitive coding (85.9 on LiveCodeBench v6), weaker at fixing real codebases (66.4 on SWE-Bench Verified, behind Qwen3.6).

Where was Kolibri trained?

On infrastructure in Germany and Finland, using 768 NVIDIA B200 GPUs, with about 20 trillion tokens of pre-training.

Kolibri is a rare thing: a genuinely new open model from a European lab, built for German and for organizations that need to know what is inside their AI, and released under a license that lets anyone use it. It reasons well, writes fast for its size, and asks a lot of memory. If you have a 64 GB Mac or a PC with 64 GB of RAM, you can run it this weekend; if you do not, a smaller model will serve you better until Kolibri's tooling matures. Jake, it turns out, does not need a server room. He needs a memory upgrade, and Ethan has already sent him the link.

📌 If you keep one line from this page

Kolibri works on only 3.46B of its 78B parameters per token: fast enough for a CPU, but it still needs all 78B in memory.

Mac with 36 GB+: MLX. PC with 64 GB+: patched llama.cpp. Ollama: not yet.

Revision note. Written October 4, 2026, the day after Kolibri's release. May your next model fit your memory and your needs.

Related