Run Liquid AI d1-3B Locally on Windows 11, Mac and Kali Linux: Install the Free Decision Model (llama.cpp, Python) and the $10M License Catch
d1-3B is Liquid AI's new open decision model, released on Hugging Face on October 5, 2026 and announced on October 7. It does not write a single word. You give it a situation and a few questions, such as "Is this customer asking for a refund?" or "Is the screen in this photo cracked?", and it returns a calibrated probability for every possible answer in one pass, in about 8 milliseconds on a gaming GPU. The surprise is its size: at 3 billion parameters, it scores 48.57 on the Decision Index, ahead of every model under 10B and level with a decision model twelve times its size. It reads photos too, and a smaller sibling, d1-omni-600M, even listens to short voice clips. Both install locally on Windows 11, Mac or Kali Linux in a few minutes and run offline on an ordinary laptop. The catch is in the license: it is free for commercial use only while your company earns under $10 million a year. And it only runs in llama.cpp builds from October 8, 2026 or later, so last week's copy fails.
Jake runs a phone repair shop, and his phone buzzes all day. Customers send messages ("is my iPhone ready?"), photos of cracked screens, voice notes recorded in the car, and the odd message that is plainly spam. Sorting them takes him an hour every evening. His friend Ethan, a developer, had already shown him chat models that could sort messages, and Jake had hated them: slow, wordy, and once a chatbot invented a category that did not exist. "I don't want a chat," Jake said. "I want a machine that says 'repair, urgent, screen cracked' and shuts up." Ethan smiled. "That's exactly what a decision model is." This page covers what Liquid AI's d1-3B and d1-omni-600M are, what the benchmarks really show, the license in plain English, how to run them with llama.cpp or Python on Windows, Mac and Kali Linux, how to write questions that work, and the small mistakes that make them look broken.
Decision models are new enough that most people have never used one. If that is you, our guide to decision models in Ollama explains the idea gently with two small models. This page stands on its own, though; everything below starts from zero.
What Liquid AI d1-3B is, in plain English
Most AI models you have used are generators: they write the next word, then the next, until an answer appears. That is wonderful for essays and terrible for sorting. If you ask a chat model "Is this message a repair request? Answer yes or no," it might say "Yes," "Yes.", "Yes, it appears to be," or occasionally something you never asked for, and your code has to cope with all of it.
A decision model works differently. You hand it two things:
- A state: the situation. A customer message, a JSON record, a photo, or a mix of them.
- Questions: one or more named questions, each with a fixed set of possible answers that you define.
It reads the state once and returns, for every question, a probability for each allowed answer. There is no writing step at all. Liquid's own description is "zero output tokens," and that is literally what the usage report says: "output_tokens": 0. Because it never generates, it cannot answer outside your options, cannot ramble, and is fast enough to call on every single message, photo or webhook.
Every question has one of three types, the same ones used by other decision models and by llama.cpp's decision endpoint:
- noul: a yes-or-no question. You get one number, the probability of "yes." The odd name is simply the format's own word for a yes-or-no question.
- choice: pick one of several named options you describe, such as "repair," "pickup question," "quote request" or "spam." You get the winner, a confidence and every option's probability.
- score: place something on a scale of 2 to 10 levels you describe in words, lowest first, such as "can wait," "today," "customer is blocked now." You get the expected level and the probability of each.
Liquid describes the answers as calibrated. That means a probability of 0.8 should be right about 80% of the time, which is what lets you set sensible thresholds: act automatically above 0.9, ask a human between 0.5 and 0.9, ignore below. One fair warning from llama.cpp's documentation: the probabilities are scaled with settings stored in the model, and they are not guaranteed to be calibrated for your data. Check a sample of your own messages before you trust a threshold.
The facts, from Liquid AI's model card:
- Parameters: 3.12 billion, built on Liquid's LFM2.5-VL-3B vision-language model.
- Vision: a 400M SigLIP2 NaFlex image encoder, so photos can be part of the state, or the whole state.
- Context: 32,768 tokens.
- Languages: the model card lists 16, including English, Spanish, French, German, Hindi, Arabic, Chinese, Japanese and Portuguese.
- License: LFM Open License 1.0, which is free below a $10 million revenue threshold. More on that below.
- Released: weights on Hugging Face October 5, 2026; official GGUF files October 6; Liquid's announcement October 7.
Liquid recommends it wherever a pipeline needs a yes/no, a pick from named options or a rating: routing and triage, moderation, intent and topic classification, extraction checks, reranking, scoring with an AI judge, guardrails for agents, and visual inspection. It is not a chat model and does not write text.
Who Liquid AI is
Liquid AI is a private company in Cambridge, Massachusetts, founded in 2023 by researchers from MIT's Computer Science and Artificial Intelligence Laboratory. Its LFM models, short for Liquid Foundation Models, are designed for phones, cars and other small devices; LFM2 and LFM2.5 chat models are already in Ollama's library. In December 2024 it raised a $250 million round led by AMD. d1 is its first family built only for decisions.
d1-3B and d1-omni-600M: the two models
| d1-3B | d1-omni-600M | |
|---|---|---|
| Parameters | 3.12B | 587M (381M trunk, 94M vision, 112M audio) |
| Built on | LFM2.5-VL-3B, a decoder | LFM2.5-Encoder-350M, an encoder |
| Reads | Text, JSON, images | Text, JSON, images, or up to 30 seconds of speech |
| Context | 32,768 tokens | 16,384 positions; with images, the text is cut to 896 tokens |
| Decision Index 0.2.1 | 48.57 | 15.95 |
| Status | The main release | "First experimental checkpoint," under active development |
In short: d1-3B is the one to build on. d1-omni-600M is a small, early research release whose party trick is audio. Liquid itself does not publish speed numbers for it yet.
d1-3B benchmarks: what the numbers really say
The headline number comes from the Decision Index 0.2.1, a public benchmark built only for decision models. Liquid scored d1-3B and d1-omni-600M with the official scorer; the other rows come from the public leaderboard. That is a fair comparison of method, with one honest caveat: Liquid's two rows are its own runs, not leaderboard submissions.
| Model | Size | Decision Index | Knowledge | Language | Retrieval | Tools | Arts |
|---|---|---|---|---|---|---|---|
| Winnow-12B | 12B | 50.02 | 33.8 | 56.0 | 54.0 | 71.0 | 30.0 |
| d1-3B | 3B | 48.57 | 23.8 | 56.4 | 52.8 | 74.5 | 36.3 |
| Decider 35B-A3B | 36B | 47.11 | 31.8 | 55.5 | 54.7 | 56.5 | 32.6 |
| JPT-9B | 9.7B | 46.89 | 31.7 | 56.7 | 44.6 | 67.0 | 28.6 |
| Decision 1.0 Lux | 9.7B | 43.49 | 30.9 | 48.0 | 50.0 | 57.2 | 26.4 |
| JPT-4B | 4.7B | 43.04 | 28.7 | 52.5 | 45.0 | 57.2 | 25.8 |
| Jet v6.2 | 4.7B | 42.60 | 28.7 | 43.9 | 48.2 | 62.9 | 27.0 |
| Decider 4B | 4.7B | 40.70 | 25.7 | 46.0 | 44.7 | 58.6 | 25.0 |
| Winnow-E4B | 8.0B | 39.89 | 22.3 | 45.1 | 43.8 | 62.5 | 22.8 |
| Decider 2B | 2.3B | 28.97 | 14.9 | 32.6 | 37.3 | 42.4 | 14.6 |
| d1-omni-600M | 587M | 15.95 | 8.3 | 12.9 | 35.0 | 15.1 | 6.8 |
Three things stand out. First, only one model in the table beats d1-3B overall, and it is four times bigger. Second, d1-3B is the best of all of them at Tools (74.5) and Arts (36.3), which matters if your decisions are things like "which tool should the agent call next?" Third, its weak spot is Knowledge, 23.8, well below the 30-plus of the bigger models. A 3B model simply knows fewer facts. If your questions need world knowledge ("Is this medicine safe with alcohol?"), a bigger decider, or a human, is the right call.
The d1-omni-600M row is not a misprint. On the full Decision Index it scores 15.95, far behind everything else. Its strength is narrower, as the next table shows.
Familiar benchmarks, asked as decisions
Liquid also turned several well-known benchmarks into decision questions. Here, the gap between the models is much smaller:
| Benchmark (what it tests) | d1-omni-600M | d1-3B | Decider 4B | Decider 2B |
|---|---|---|---|---|
| SQuAD 2.0 (does the text answer the question?) | 74.0 | 85.3 | 76.0 | 67.7 |
| Civil Comments (toxicity) | 95.8 | 93.0 | 92.8 | 93.6 |
| MASSIVE intent (what does the user want?) | 86.1 | 87.3 | 88.3 | 81.1 |
| PubMedQA (medical yes/no/maybe) | 61.3 | 66.0 | 63.3 | 65.7 |
| BoolQ (yes/no from a passage) | 77.7 | 86.7 | 89.0 | 87.3 |
| XNLI (does A imply B, many languages) | 74.7 | 85.0 | 88.6 | 85.0 |
| PAWS-X (same meaning, reworded?) | 79.5 | 76.9 | 69.8 | 59.5 |
| Mean | 78.4 | 82.9 | 81.1 | 77.1 |
Read that table as a shopkeeper would. For everyday sorting, such as toxicity, intent and "same meaning?", even the tiny 587M model is in the same league as models several times its size, and it tops two rows. For questions that need careful reading, such as "does this passage really answer the question?", d1-3B pulls ahead. d1-3B also scores 71.8 on DecisionBench (all 23,900 English rows) and 69.3 on Fast Decisions.
Does it really look at the photo?
On eleven public image benchmarks asked as decisions, d1-3B averages 74.1, slightly above the 73.9 of the vision model it was built from. The most reassuring test is the one Liquid ran on purpose: with the images removed, the same questions score 45.1. In other words, the answers really come from the pictures, not from guessing based on the question. Its best image result is 88.5 on POPE, a test of whether a model claims to see objects that are not there.
How fast is d1-3B? Real latency numbers
Speed is the whole reason to use a decision model, so here are Liquid's measurements for d1-3B, warm calls, one request at a time. "64 states, packed" is throughput when many requests are batched together.
| Hardware | 1 question | 3 questions, one pass | 3.4K-token state | 384 px image | 64 states, packed |
|---|---|---|---|---|---|
| NVIDIA RTX 4090 | 8 ms | 21 ms | 102 ms | 17 ms | 475 per second |
| AMD MI325X | 9 ms | 14 ms | 44 ms | 18 ms | 1,106 per second |
| Apple M5 Pro | 30 ms | 41 ms | 640 ms | 62 ms | 78 per second |
| NVIDIA Jetson AGX Thor | 16 ms | 20 ms | 220 ms | 35 ms | 262 per second |
| NVIDIA Jetson AGX Orin 64 GB | 26 ms | 35 ms | 560 ms | 83 ms | 110 per second |
| NVIDIA Jetson Orin Nano | 50 ms | 73 ms | 1,640 ms | 202 ms | 38 per second |
Two useful patterns hide in there. Asking three questions at once costs only a little more than asking one, about 1.3 times the time in Liquid's words, because the state is read once and shared by every question. So always bundle your questions into one call. And long text is the expensive part: a 3,400-token state takes several times longer than a photo. Keep states to what the decision needs.
The RTX 4090 figure of 8 ms uses a compiled mode in Python (model.compile(mode="reduce-overhead")); without it, a single question takes 16 ms. These are Liquid's numbers in its own Python code. llama.cpp on an ordinary laptop processor will be slower, but for one message at a time "slower" usually still means well under a second.
The license catch: free under $10 million, and what that means
This is the part most launch coverage skipped, and it is the one a business owner needs first. Most of the decision models we have covered use Apache 2.0, which lets anyone use them commercially with no conditions. d1 uses the LFM Open License 1.0, which is different in one important way.
In plain English:
- Anyone may download, use, modify and share it, including fine-tuned versions, with a free copyright and patent license.
- Commercial use is free only while your company stays under the "Threshold": annual revenue of $10,000,000 or more. The license defines commercial use broadly, as any use "for direct or indirect commercial advantage or monetary compensation."
- Above the threshold, commercial use is simply not licensed under this agreement. A larger company needs a separate commercial agreement with Liquid AI.
- Qualified non-profits, such as charities and educational bodies with 501(c)(3) status or the foreign equivalent, are exempt from the threshold when the use is non-commercial or research.
- The usual Apache-style terms apply on top: keep the license and notices when you redistribute, and the patent license ends if you sue a contributor over patents in the model.
For Jake's shop, a solo developer, a startup or a small agency, this is effectively free. For a company over $10 million in revenue, it is a model you can test but not ship without talking to Liquid. If you are building a product you hope will grow past that line, factor it in now rather than at your Series B. If you need an Apache-licensed decider instead, our guides to AWS's Strands Decider 2B and the open llama.cpp decision models cover the options.
This is a summary, not legal advice. The full license is a short file next to the model on Hugging Face, and it is worth five minutes of your time before you deploy anything commercial.
d1-3B files, sizes and the hardware you need
Liquid publishes official GGUF files for llama.cpp. Images need a second, separate file called an mmproj, short for multimodal projector; text-only use does not.
| File | d1-3B | d1-omni-600M |
|---|---|---|
| Q4_K_M model | 1.67 GB | not published |
| Q8_0 model (used in Liquid's examples) | 2.87 GB | 407 MB |
| F16 or BF16 model | 5.4 GB | 764 MB |
| Q8_0 mmproj (images, and audio on omni) | 580 MB | 263 MB |
| F16 or BF16 mmproj | about 850 MB | about 440 MB |
Translated into machines: d1-3B at Q8_0 with images needs about 3.5 GB of memory plus a little working space, so any laptop with 8 GB of RAM runs it, on the processor alone if necessary. A graphics card with 4 GB or more makes it fast. d1-omni-600M fits in well under 1 GB and runs on almost anything, including small boards. For a sense of where these sit against chat models, our laptop guide for local AI has the tiers; d1 lives comfortably at the bottom of every one of them.
Which file? Liquid's own command uses Q8_0, and for a decision model, where a small shift in probabilities can flip a borderline answer, that is the sensible default. Use Q4_K_M only when memory is genuinely tight, and check a few borderline cases against Q8_0 when you do.
How to install d1-3B locally with llama.cpp (Windows 11, Mac, Linux and Kali)
This is the local install most people want: one download, one command, and the model runs offline on your own machine. llama.cpp is the engine underneath Ollama and LM Studio, and its small web server, llama-server, has a dedicated decision endpoint called /v1/systemone. The same endpoint serves other open decision models, so code you write for d1 works with them too. Our llama.cpp decision-model guide explains the endpoint in depth; here is the short path for d1.
Step 1: get a new enough llama.cpp
This is the step that decides everything. Support for d1-3B was added to llama.cpp on October 8, 2026, in build b11483, and d1-omni-600M followed the same day in b11488. Any copy older than that, even one you downloaded last week for another decision model, will not load d1 correctly.
- Windows: download the latest release zip from the llama.cpp releases page on GitHub (pick the build for your graphics card, or the CPU one), or run
winget install llama.cppand thenwinget upgrade llama.cpp. - Mac:
brew install llama.cpp, orbrew upgrade llama.cppif you already have it. - Linux and Kali: download a prebuilt Linux release from the same page, or build from source with CMake for the newest version.
Then check: llama-server --version. The build number should be 11488 or higher to cover both models.
Step 2: start the model
llama-server -hf LiquidAI/d1-3B-GGUF:Q8_0
The -hf flag downloads the model from Hugging Face on first run, together with its image projector, and caches it, so later starts take seconds. The server listens on http://127.0.0.1:8080. For the omni model, Liquid adds two flags, because a whole question must fit in one processing batch:
llama-server -hf LiquidAI/d1-omni-600M-GGUF:Q8_0 -b 4096 -ub 4096
4096 covers any image and states of a few thousand tokens; raise both up to 16384 for longer states.
Step 3: ask your first questions
This is Liquid's own example, a customer message with three questions answered in one pass:
curl http://127.0.0.1:8080/v1/systemone -H "Content-Type: application/json" -d '{
"state": "I was charged twice this month, please refund one of them.",
"questions": {
"refund": {"type": "noul", "instructions": "Is the customer asking for a refund?"},
"team": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "Charges, refunds, invoices",
"technical": "App or site faults",
"fraud": "Suspected unauthorised use"}},
"urgency": {"type": "score", "instructions": "How urgent is this?",
"criteria": ["Can wait", "Today", "Blocking the customer now"]}
}
}'
The reply contains an answers object with your three names as keys: a probability of "yes" for refund, the winning team with a confidence and every team's probability for team, and an expected level with a probability per level for urgency. On Windows, PowerShell treats quotes differently, so the easiest route there is the Python example further down, or saving the JSON to a file and sending it with curl.exe ... -d @request.json.
Step 4: add a photo
Images go in an images list as base64 data URLs. The state can be null when the photo is the whole story:
curl http://127.0.0.1:8080/v1/systemone -H "Content-Type: application/json" -d @- <<JSON
{
"state": "Photo sent by a customer with a repair request.",
"images": ["data:image/jpeg;base64,$(base64 < phone.jpg | tr -d '\n')"],
"questions": {
"cracked": {"type": "noul", "instructions": "Is the phone screen cracked?"},
"device": {"type": "choice", "instructions": "What kind of device is this?",
"criteria": {"phone": "A smartphone", "tablet": "A tablet",
"laptop": "A laptop", "other": "Something else"}}
}
}
JSON
That command works in bash and zsh on Mac, Linux and Kali, including Kali's default zsh. On Windows, use the Python route below.
Step 5: call it from Python
No special library is needed; the endpoint is plain JSON over HTTP:
import base64, json, urllib.request
URL = "http://127.0.0.1:8080/v1/systemone"
def decide(state, questions, image_path=None):
body = {"state": state, "questions": questions}
if image_path:
with open(image_path, "rb") as f:
body["images"] = ["data:image/jpeg;base64," + base64.b64encode(f.read()).decode()]
req = urllib.request.Request(URL, data=json.dumps(body).encode(),
headers={"Content-Type": "application/json"})
with urllib.request.urlopen(req) as r:
return json.load(r)["answers"]
print(decide("Is my iPhone ready? Dropped it off Monday.",
{"intent": {"type": "choice", "instructions": "What does the customer want?",
"criteria": {"status": "Asking whether a repair is finished",
"quote": "Asking for a price",
"booking": "Wants to book a repair",
"spam": "Not a real customer"}}}))
Run d1-3B or d1-omni-600M in Python with Transformers
If you would rather skip llama.cpp, both models run in Python through Hugging Face Transformers. d1-3B needs transformers 5.14 or later; d1-omni-600M needs 5.15 or later. Both ship their own code, so they load with trust_remote_code=True. Only do that for repositories you trust; it runs the publisher's Python on your machine.
Create a virtual environment first. On Kali and other recent Linux systems, a system-wide pip install is refused with "externally-managed-environment," and our guide to that error explains why. Then install, keeping the quotes:
python3 -m venv ~/d1
source ~/d1/bin/activate
pip install "transformers>=5.15" torch torchvision pillow soundfile
On Windows, the activate line is d1\Scripts\activate. Those quotes around "transformers>=5.15" are not decoration: without them, every shell reads > as "send output to a file," installs whatever version it likes, and leaves a stray file named =5.15 in your folder.
import torch
from transformers import AutoModel
device = "cuda" if torch.cuda.is_available() else "mps" if torch.backends.mps.is_available() else "cpu"
dtype = torch.float32 if device == "cpu" else torch.bfloat16
model = AutoModel.from_pretrained("LiquidAI/d1-3B", trust_remote_code=True, dtype=dtype).to(device)
questions = {
"repair": {"type": "noul", "instructions": "Is this a request to repair a device?"},
"urgency": {"type": "score", "instructions": "How urgent is this for the customer?",
"criteria": ["Can wait", "This week", "Today", "Phone is unusable now"]},
}
print(model.system_one("My screen went black and I need my phone for work tomorrow.", questions))
Three calls are available: system_one(state, questions, images=None) for one request, system_one_batch([...]) to pack many requests together with no padding, and, on the omni model, probabilities(...) for the raw distributions. Each call returns {"answers": {...}, "usage": {"input_tokens": n, "output_tokens": 0}}.
The number types matter, and they differ between the two models:
- d1-3B: float32 on a processor, bfloat16 on GPUs and Apple Silicon, as in Liquid's example.
- d1-omni-600M: float32 on a processor, float16 on GPUs. Liquid checked that float16 gave the same top answer as float32 on every one of 243 text, 214 image and 416 audio test rows, while bfloat16 changed the top answer on 0.8% of text and 1.7% of audio rows. Small, but in a decision model a changed answer is a wrong answer.
How to write questions that d1-3B answers well
A decision model is only as good as its questions, and this is where most of the quality comes from. The question schema has three parts: type, instructions and, for choice and score, criteria. Here is what works:
- Keep instructions short and literal. "Is the customer unhappy?" beats a paragraph about tone and sentiment.
- Describe every option in its criteria.
"status": "Asking whether a repair is finished"is far clearer than a bare"status". The model reads the description, not your variable name. - Make options cover everything, with an exit. Add "other" or "spam" so the model is not forced to pick a wrong real option.
- Keep options from overlapping. If "quote" and "price question" mean the same thing, merge them, or the probability splits between them and neither wins clearly.
- Order score levels from lowest to highest. The model expects the scale to run upward.
- Define both sides of a hard yes/no. A noul question accepts optional criteria such as
{"true": "Visible crack or shattered glass", "false": "Screen intact, even if dirty"}, which removes the gray zone. - Bundle related questions. Three questions in one call cost about 1.3 times one question.
- Put only what matters in the state. Long states are the slow part. Strip email signatures and quoted reply chains.
Then set thresholds from the probabilities instead of trusting every top answer. Ethan's starting rule for Jake: act automatically above 0.9, flag for a human between 0.6 and 0.9, and treat anything below 0.6 as "not sure." Then check them against a week of real messages and adjust, because a threshold is only as good as the sample it was tested on.
A real example: Jake's message and photo triage
Here is the small script Ethan built for Jake, in plain Python with the llama.cpp server running on the shop's laptop. It reads each new customer message, with its photo if there is one, asks four questions in one call, and files the message into a folder for Jake's evening, or, when the model is sure, sends an automatic reply.
QUESTIONS = {
"intent": {"type": "choice", "instructions": "What does the customer want?",
"criteria": {"status": "Asking whether their repair is finished",
"quote": "Asking how much a repair costs",
"booking": "Wants to book or drop off a device",
"complaint": "Unhappy with a repair or the service",
"spam": "Not a real customer message"}},
"urgency": {"type": "score", "instructions": "How urgent is this for the customer?",
"criteria": ["Can wait", "This week", "Today", "Device unusable right now"]},
"upset": {"type": "noul", "instructions": "Is the customer upset or angry?"},
}
PHOTO_QUESTION = {"cracked": {"type": "noul", "instructions": "Is the device screen cracked?",
"criteria": {"true": "Visible crack or shattered glass",
"false": "Screen intact, even if dirty"}}}
def triage(text, photo=None):
qs = dict(QUESTIONS, **(PHOTO_QUESTION if photo else {}))
a = decide(text, qs, photo) # decide() from the Python example above
intent = a["intent"]
if intent["choice"] == "spam" and intent["confidence"] > 0.9:
return "spam"
if a["upset"]["noul"] > 0.7:
return "call-back-today" # complaints never get an automatic reply
if intent["choice"] == "status" and intent["confidence"] > 0.9:
return "auto-reply-status"
return "evening-pile"
Notice the design choices. Upset customers always reach a human, whatever else the model thinks. Automatic replies only happen for the safest category, "is it ready?", and only above 0.9 confidence. Everything else lands in Jake's evening pile, now sorted. And the model never writes a word to a customer; the replies are Jake's own templates.
The first week, Jake checked every decision by hand, as Ethan insisted. He found the model was cautious in a good way: when it was unsure, it said so with a low confidence, and those were exactly the messages a human should read. "It's like a new shop assistant who asks when in doubt," Jake said. "Which is more than I can say for the last one."
That first week is worth copying exactly. Here is the rollout order that keeps a new decision model from embarrassing you in front of customers:
- Shadow mode for a week. Let the model decide every message, but act on none of its answers. Log what it said next to what you actually did.
- Compare on real cases. Go through the log and mark each decision right or wrong. Look hardest at the wrong ones with high confidence; those show a question that needs rewording.
- Fix the questions, not the model. Add missing options, merge overlapping ones, and define both sides of any yes/no that keeps going wrong.
- Set thresholds from your log. Choose the confidence above which the model was almost never wrong on your messages, not on a benchmark.
- Automate the safest category first. For Jake that was "is my repair ready?" Everything else stays with a human until the log says otherwise.
- Re-check every month. Customers change how they write, and new kinds of messages arrive. Ten minutes with the log keeps the thresholds honest.
d1-omni-600M: decisions about voice notes
The smaller model's unusual skill is audio. It takes up to 30 seconds of 16 kHz mono speech, or text with images, never both in one request, and answers the same three question types. In llama.cpp, audio goes in a files list:
curl http://127.0.0.1:8080/v1/systemone -H "Content-Type: application/json" -d @- <<JSON
{
"state": "Voice note from a customer.",
"files": ["data:audio/wav;base64,$(base64 < note.wav | tr -d '\n')"],
"questions": {
"kind": {"type": "choice", "instructions": "What kind of utterance is this?",
"criteria": {"request": "A request to do something",
"question": "A question asking for information",
"other": "Something else"}}
}
}
JSON
Read Liquid's own warning before relying on it. The audio skills were trained on requests between an English speaker and an assistant: what kind of utterance it is, its topic and what the speaker wants. It is not a general audio classifier, it will not identify a song or a dog bark reliably, and clips longer than 30 seconds are cut. If you convert phone voice notes first, make them mono 16 kHz WAV, for example with ffmpeg -i note.m4a -ac 1 -ar 16000 note.wav.
For text and images, the omni model has one more limit worth knowing: when a request includes images, the state and question text is cut to 896 tokens, as it was trained. Keep image-request text short.
Why d1-3B looks broken: the common mistakes
Decision models fail quietly. They rarely crash; they just return answers that look plausible and are wrong. Here is the list, most common first.
A llama.cpp build from before October 8
d1-3B needs build b11483 or later, and d1-omni-600M needs b11488 or later. An older build may refuse the model or answer badly. Run llama-server --version and update first; this one check saves most first-day frustration.
Error 501 when you send a photo
llama-server returns 501 for an image request when the model has no multimodal projector loaded. With -hf, the projector normally comes along automatically; if you load a local file with -m, add --mmproj with the path to the mmproj-d1-3B-Q8_0.gguf file. A 501 can also mean the loaded model is not a decision model at all.
Running it in Ollama or LM Studio
d1 is not in Ollama's model library; Ollama carries Liquid's LFM2 and LFM2.5 chat models, which are different models. LM Studio's chat screen is built for generating text, which d1 does not do. The documented route is llama-server's /v1/systemone, or Python.
Asking it to write something
"Summarize this message" or "draft a reply" cannot work: a decision model has no writing step. Turn every task into a question with fixed answers, and let your own templates or a separate chat model do any writing.
Unquoted version numbers in pip
pip install transformers>=5.15 without quotes writes to a file called =5.15 and installs the wrong version. Always quote: pip install "transformers>=5.15".
bfloat16 on the omni model
For d1-omni-600M, use float16 on GPUs and float32 on processors. bfloat16 flipped the top answer on up to 1.7% of Liquid's audio test rows.
Images and audio in the same request
The omni model takes images or one audio clip per request, not both. In Python, passing both raises a ValueError. Split them into two calls.
Long text with images on the omni model
With images, the omni model cuts the state and questions to 896 tokens. A long email plus a photo loses most of the email silently. Keep the text short, or use d1-3B for long text.
Options that overlap or do not cover everything
If two options mean the same thing, the probability splits and neither wins. If no option fits, the model still has to pick one. Merge duplicates and always include an "other" option.
Batch size too small on the omni model
The omni model reads each question in one batch, so the batch must hold it. Start the server with -b 4096 -ub 4096, and go up to 16384 for long states, as Liquid's GGUF page says.
Forgetting the license threshold
A prototype at a company with $10 million or more in annual revenue is fine for testing, but shipping it commercially is not covered by the free license. Check before launch, not after.
d1-3B vs other decision models and vs chat models
We have covered most of the open decision models as they arrived, so here is where d1 fits among them, without mixing up anyone's benchmark claims:
| Option | Where it runs | License | Choose it when |
|---|---|---|---|
| d1-3B | llama.cpp, Python | LFM Open 1.0 (free under $10M revenue) | You want the strongest small decider, with photos, on modest hardware |
| d1-omni-600M | llama.cpp, Python | LFM Open 1.0 | You need short English voice commands, or the tiniest footprint |
| Nimble and Tev1 in Ollama | Ollama | Apache 2.0 | You already use Ollama and want the simplest start |
| Strands Decider 2B (AWS) | Its own package and server | Apache 2.0 | You need a permissive license at any company size |
| Cloudflare Clef | Cloudflare Workers AI, also local | Apache 2.0 | You want a hosted option next to your Cloudflare apps |
| A chat model asked to classify | Anywhere | Varies | The decision needs explanation, long reasoning or world knowledge |
Because they share the /v1/systemone request shape, switching between d1 and the other llama.cpp deciders is mostly a matter of changing which model the server loads. That makes an honest bake-off cheap. Here is a fair way to run one:
- Collect fifty real cases with the answer you know is right for each, including a few hard ones.
- Write the questions once and use exactly the same questions for every model.
- Run every case through each model, swapping only the model the server loads.
- Count correct top answers, then look at confidence: a model that is wrong but unsure is safer than one that is wrong and certain.
- Check speed and license last. If two models tie on accuracy, the faster one with the license that fits your company wins.
Keep the winner, and keep the fifty cases too; they become your test set the next time a new decider appears. Our guides to Cloudflare Clef and the bigger open Jev models cover the alternatives in depth.
And against a chat model doing the same job? The decision model wins on speed, cost, consistency and never answering outside your options. The chat model wins when you need a reason, a summary or knowledge the small model lacks. Many real systems use both: the decider routes and gates thousands of items cheaply, and a chat model handles the few that need words.
Running d1 on AWS or a small server
d1 is small enough that a server is optional; the shop laptop handled Jake's whole day. If you want it running for a website or an app, the simplest pattern is a small always-on machine running llama-server, kept private: listen on the private network only, and put authentication in front of it if anything outside must reach it. llama-server listens on 127.0.0.1 by default, which is exactly right until you decide otherwise.
On AWS, that means a modest EC2 instance in your VPC; a GPU helps with throughput but is not required for message-by-message traffic. For heavy batch work, such as scoring a million product photos overnight, a GPU instance and Python's system_one_batch packing are the efficient route. Liquid's models are not in Amazon Bedrock's serverless catalog, so this is self-hosting either way. Our explainer on Hugging Face models on AWS walks through the deployment options and their costs, and the $10 million license rule applies wherever the model runs.
Liquid AI d1-3B: frequently asked questions
What is Liquid AI d1-3B?
d1-3B is an open decision model from Liquid AI, released October 5, 2026. Given a state of text, JSON or images and a set of yes/no, choice or rating questions, it returns a probability for every answer in one pass, with zero output tokens. It has 3.12B parameters.
What is a decision model?
A decision model answers questions with fixed options instead of writing text. You define the questions and allowed answers; it returns a probability for each. That makes it fast, cheap and unable to answer outside your options, ideal for routing, moderation and triage.
Is Liquid AI d1-3B free?
It is free to download and use. Commercial use is free while your company's annual revenue is under $10 million, under the LFM Open License 1.0. Larger companies need a separate agreement with Liquid AI for commercial use.
How do I run d1-3B locally?
Install a llama.cpp build b11483 or newer, run llama-server -hf LiquidAI/d1-3B-GGUF:Q8_0, then POST a state and questions to http://127.0.0.1:8080/v1/systemone. Or load LiquidAI/d1-3B in Python with transformers 5.14 or later.
Can I run d1-3B in Ollama?
Not today. d1 is not in Ollama's library; Ollama has Liquid's LFM2 and LFM2.5 chat models, which are different. Use llama-server's /v1/systemone endpoint or Python Transformers.
How much RAM does d1-3B need?
The Q8_0 model is 2.87 GB and its image projector about 580 MB, so around 4 GB of memory covers it. Any laptop with 8 GB of RAM runs it. The Q4_K_M file is 1.67 GB for tighter machines.
How fast is d1-3B?
Liquid measured 8 ms per question on an RTX 4090, 30 ms on an Apple M5 Pro and 50 ms on a Jetson Orin Nano. Three questions in one call take about 1.3 times one question. A 3,400-token state takes 102 ms on the RTX 4090.
How good is d1-3B?
It scores 48.57 on Decision Index 0.2.1, best among models under 10B and level with Decider 35B-A3B, which is about 12 times bigger. Its weak area is world knowledge, 23.8 versus over 30 for larger models.
What is d1-omni-600M?
The smaller, experimental sibling: 587M parameters, reading text, images or up to 30 seconds of English speech. It scores 15.95 on the Decision Index but holds its own on everyday tasks such as toxicity and intent, with a 78.4 average on Liquid's benchmark set.
Can d1 understand voice notes?
d1-omni-600M can, within limits. It takes one 16 kHz mono clip of up to 30 seconds, and its audio skills were trained on English requests to an assistant: utterance type, topic and intent. It is not a general sound classifier.
What are noul, choice and score questions?
noul is yes or no and returns the probability of yes. choice picks one of your named options and returns each option's probability plus a confidence. score places the state on a 2 to 10 level scale you describe, lowest first.
Are d1 probabilities calibrated?
Liquid calibrates them, and they come scaled by settings stored in the model. llama.cpp's documentation notes they are not guaranteed to be calibrated for your data, so test your thresholds on a sample of your own cases.
Why does llama-server return error 501 for images?
No multimodal projector is loaded, or the model is not a decision model. With -hf the projector usually downloads automatically; with -m, add --mmproj pointing to the mmproj-d1-3B GGUF file.
How many options can a choice question have?
llama.cpp allows up to 255 options per choice question for both d1 models. In practice, fewer clear and non-overlapping options give more confident answers than a long list.
How do I install Liquid AI d1-3B locally on Windows 11?
Install llama.cpp with winget install llama.cpp (build 11483 or newer), open a terminal and run llama-server -hf LiquidAI/d1-3B-GGUF:Q8_0. The 2.9 GB model and its image projector download once; then send questions to http://127.0.0.1:8080/v1/systemone. No graphics card is required.
Does d1-3B work offline?
Yes. After the first download, llama-server and the Python package run with no internet connection. Nothing you send to the model leaves your computer.
Is d1-3B on Amazon Bedrock?
No. Self-host it on an EC2 instance or any server with llama-server or Python. It is small enough that a GPU is optional for message-by-message work.
Who makes d1?
Liquid AI, a company founded in 2023 by researchers from MIT's Computer Science and Artificial Intelligence Laboratory, based in Cambridge, Massachusetts. It also makes the LFM2 and LFM2.5 model families.
Can I use d1-3B on Kali Linux?
Yes. Use a recent llama.cpp release binary or build it with CMake, then run llama-server. For Python, create a virtual environment first, because Kali refuses system-wide pip installs, and quote "transformers>=5.14".
A week later, Jake's evening hour had shrunk to fifteen minutes. The spam was gone before he saw it, the "is it ready?" messages answered themselves, and the cracked-screen photos arrived already tagged. The upset customers still came to him first, which was the rule he liked best. "It doesn't talk," he told Ethan, "and it doesn't make things up." That is the quiet promise of a decision model, and d1-3B delivers it at a size that fits on any laptop in the shop.
If you keep one line from this page
Ask small models fixed questions, keep a human on the upset ones, and check the license before you grow.
Update llama.cpp to b11488 or later, describe every option, and test thresholds on your own data.
Revision note. Written October 9, 2026, the day after llama.cpp added both d1 models. If an hour of sorting messages is eating your evenings too, this is a gentle place to start getting it back.
