AWS Strands Decider 2B: The Free Decision Model That Runs on Your PC (Setup, vs Jev and Clef)

Logeshwaran
—

AWS Strands Decider 2B is a free, open-source decision model from AWS's Strands Labs team, released at the start of October 2026. It does not write text. You give it some text (the "state") and typed questions, such as yes/no, pick one option, or rate on a scale, and it returns the answer with a calibrated confidence, in about 115 milliseconds on a gaming GPU and about 150 milliseconds on an M3 Mac. It has 1.9 billion parameters, is built on Qwen3.5-2B, uses the Apache 2.0 license, needs no AWS account, and installs with pip install strands-decider. Its server speaks the same /v1/systemone request format as TypeSafe's Jev, Ollama's decision models and llama.cpp, so it can replace them without code changes. It scores 0.723 on JevBench, first among true 2B models. Setup on Windows, Mac and Kali, the honest comparison and the limits are below.

Jake's online booking form had a small problem that cost him an hour every morning: customers typed whatever they liked into the "what's wrong with your phone" box, and someone had to read each one and send it to the right person, screen repairs to one bench, battery jobs to another, warranty claims to the counter. Ethan had already tried a chatbot for it, and it worked, but it was slow, it sometimes wrote a paragraph instead of a team name, and every message cost a few tokens. When AWS released Strands Decider 2B, he tried it on a hundred real messages that evening. It never wrote a word, just a choice and a confidence number, it answered each one in about a tenth of a second on Jake's old gaming PC, and when its confidence was low, the message went to a human instead. That is what decision models are for, and this one is the smallest serious open option so far.

⚡ Quick Answer

• What it is → a 1.9B model that picks answers, never writes. Specs.

• Install → pip install strands-decider in a virtual environment. Setup.

• Use it → CLI, Python, or the /v1/systemone server. Server.

• Compared → vs Jev, Clef, Kev and friends. Table.

• Tested → works on a 2016 laptop, about 5 s per decision on CPU. Our test.

Not on Bedrock, no AWS account needed, Apache 2.0. AWS questions.

🧭 NEW HERE? READ THESE FIRST

New to decision models? These five pages pair with this one:

 Bookmark this; small decisions do not need a chatbot.

First, what is a decision model?

Most AI models people know are generative: you ask, and they write an answer word by word. That is perfect for writing, and wasteful for decisions. If all you need is "billing, sales or retail?", a chatbot still has to generate text, which takes time, costs tokens, and sometimes produces a sentence instead of one of your options.

A decision model skips the writing. It reads the text once and scores your options directly, returning the winner and a probability. Because it can only ever answer with one of the options you gave it, it cannot ramble, invent a new category, or need its answer parsed. The idea took off in 2026 with TypeSafe's hosted Jev model and its "System One" request format, and open models followed: Bespoke's Nimble and Together's Tev1 in Ollama, Kev and OpenJev in llama.cpp, and Cloudflare's Clef. Our plain-English guide to decision models covers the idea in depth.

Where do they fit? Anywhere a program needs a quick judgment on text: routing support tickets, flagging urgent messages, checking whether an AI agent should be allowed to run a tool, deciding if a document answers a question, rating sentiment. They are the reflexes; a full language model stays the brain for the jobs that need writing.

What Strands Decider 2B is: the specs

DetailStrands Decider 2B
MakerAWS Strands Labs (the team behind the Strands Agents SDK)
ReleasedStart of October 2026 (weights posted September 30)
Model nameStrandsAgents/strands-decider-2B-hobson-v19 (Hugging Face)
Size1.9 billion parameters
Built fromQwen3.5-2B-Base, fine-tuned with LoRA (rank 16), text head replaced by a small pointer head (about 1M parameters)
Question typesYes/no ("noul"), choice (one of N), score (ordered scale)
Context4,096 tokens
SpeedAbout 115 ms median on an RTX 3090; about 153 ms on an M3 Pro for short inputs
Runs onNVIDIA GPU (CUDA), Apple Silicon (MPS or MLX), or CPU
LicenseApache 2.0: free, commercial use allowed
Codegithub.com/strands-labs/strands-decider (package, server, training recipe)

How it decides without writing a word

Inside, Strands Decider is a normal small language model with one part swapped out. A language model ends in a "head" that turns its understanding of the text into the next word. AWS removed that head and attached a tiny pointer head instead, which looks at the model's understanding of the text and of each option you supplied, and points at the best match. All your options are scored together in a single pass, and the scores become probabilities. There is no word-by-word generation loop at all, which is why it is fast and why it can only ever answer with one of your options.

The answers come with calibrated confidence. Calibration means a 0.8 really is right about 80% of the time. AWS reports an expected calibration error of 0.052 and a Brier score of 0.342 on its benchmark, which in plain terms means the confidence numbers are meaningful enough to act on: auto-route the confident answers, and send the uncertain ones to a person.

The three question types, with examples

  • Yes/no, which the project calls noul: "Does this message convey urgency?" The answer is a probability between 0 and 1.
  • Choice: "Which team should handle this?" with options such as billing, sales, retail. The answer is the winning option, with probabilities for each.
  • Score: "How frustrated is the writer?" on an ordered scale such as calm, frustrated, very upset. The answer is a position on that scale.

You can ask several questions about the same text in one request, and each gets its own answer. Phrase them plainly and specifically: the model card notes that rephrasing a question can change the answer, so test your exact wording on real examples before relying on it.

How good is it? The honest numbers

TestAccuracyItems
JevBench public set0.723167 of 231
Held-out short tasks0.6416,000
Board-game reasoning0.822900
ContractNLI (contract clauses)0.8721,026
HotpotQA (held out)0.717959
MuSiQue (multi-step questions)0.8841,199

On JevBench, the public benchmark for this kind of model, 0.723 makes it the best true 2B model and third overall out of 33 models tested, which is remarkable for its size. These are AWS's own numbers, though, with the evaluation code published so anyone can rerun them. Two notes keep them in proportion: the "held-out short tasks" score (0.641) is closer to what you will see on unfamiliar real-world tasks, and some of the datasets above were part of its training mix, which flatters those scores.

Strands Decider vs Jev, Clef, Kev and the rest

ModelSizeWhere it runsLicenseBest for
Strands Decider 2B1.9BYour machine, via its own package and serverApache 2.0Fast local gating, agent guardrails, small hardware
Jev (TypeSafe)UndisclosedHosted API onlyClosedTop accuracy with no hardware
Clef / Clef-flash (Cloudflare)27B / 9BWorkers AI, Ollama, llama.cppApache 2.0Questions about images; more capacity
Kev (0.8B, 4B, 9B)0.8B to 9Bllama.cppApache 2.0Choosing a size to fit your hardware
OpenJev27Bllama.cppNon-commercialResearch and personal use
Nimble, Tev1SmallOllamaOpenPeople who already live in Ollama

The case for Strands Decider is the combination: open and commercial-friendly, small enough for almost any computer, fast, calibrated, with the full training recipe published so you can train your own version on your own labels. The case against: it is not inside Ollama or llama.cpp's official decision collection yet, and for questions about images or very long documents, larger models such as Clef do better. Our guides to Cloudflare Clef and OpenJev and Kev in llama.cpp cover those routes.

Install Strands Decider (Windows, Mac, Linux and Kali)

It is a Python package, so the setup is the same everywhere: a virtual environment, then one pip command. The first run downloads about 4.4 GB of model files from Hugging Face, so allow time for that once (on our test machine, about 15 minutes in total).

Windows

  1. Install Python 3 from python.org if you do not have it (check "Add python.exe to PATH" in the installer).
  2. Open PowerShell in a folder for the project and create an environment: python -m venv venv
  3. Activate it: .\venv\Scripts\Activate.ps1 (if scripts are blocked, run Set-ExecutionPolicy -Scope CurrentUser RemoteSigned once).
  4. Install: pip install strands-decider
  5. With an NVIDIA card, install the CUDA build of PyTorch from pytorch.org's selector so it uses the GPU; without one, it runs on the CPU, more slowly.

Mac (Apple Silicon)

  1. In Terminal: python3 -m venv venv && source venv/bin/activate
  2. pip install strands-decider
  3. For extra speed, install the MLX extra as the project's README describes; it reports roughly 1.4 to 1.6 times faster decisions on Apple Silicon.

Linux and Kali

python3 -m venv ~/venvs/decider
source ~/venvs/decider/bin/activate
pip install strands-decider

On Kali, Debian and Ubuntu, the virtual environment is required: a plain pip install outside one is blocked with "externally-managed-environment". Our guide to pip, venv and pipx on Kali explains why, from the basics. If you plan to use Strands Decider only as a command-line tool, pipx install strands-decider works too.

Your first decision from the command line

The package includes a command that asks questions straight from the terminal. This example is the one from the model card:

strands-decider ask StrandsAgents/strands-decider-2B-hobson-v19 \
  --state "Help! My payouts have been failing for 3 days!" \
  --choice "Which team should handle this?=billing,sales,retail" \
  --noul "Does this convey urgency?" \
  --score "How frustrated is the writer?=calm,frustrated,depressed"

The --state is the text to judge. Each question flag adds one question: --choice and --score take the question, an equals sign, and the options separated by commas; --noul takes a yes/no question. You get back the chosen team, the urgency probability and the frustration score, each with its confidence. (On Windows PowerShell, replace the backslash line breaks with backticks, or put the command on one line.)

Run it as a server: the /v1/systemone API

For apps, run the built-in server:

strands-decider serve StrandsAgents/strands-decider-2B-hobson-v19 --port 8000

It listens at /v1/systemone, the same endpoint name and request shape used by Jev, Ollama's decision models and llama.cpp's decision server. A request sends the state and named questions:

curl http://127.0.0.1:8000/v1/systemone -H "Content-Type: application/json" -d '{
  "state": "Help! My payouts have been failing for 3 days!",
  "questions": {
    "urgent": {"type": "noul", "instructions": "Does this convey urgency?"}
  }
}'

The reply contains an answers object with one entry per question (for a yes/no question, a probability such as 0.83), plus token usage and the measured latency in milliseconds. Because the format matches the other decision servers, an app written for any of them can point at Strands Decider by changing only the address, which makes it easy to compare models on your own data.

Keep it private. The bundled server has no login or authentication. Leave it on your own machine (127.0.0.1), or put it behind something that checks who is calling, before exposing it on a network.

Use it from Python

One warning first: the project's README shows from strands_decider import DeciderModel, but in the released package (version 0.1.0) that name does not exist, and the import fails. We checked the package itself: the code below is what its own command-line tool uses, and we ran it on our test laptop.

from strands_decider.infer import load_engine
from strands_decider.schema import NoulQuestion, ChoiceQuestion, ScoreQuestion

engine = load_engine("StrandsAgents/strands-decider-2B-hobson-v19", device="cpu")

questions = {
    "urgent": NoulQuestion(instructions="Does this convey urgency?"),
    "team": ChoiceQuestion(instructions="Which team should handle this?",
                           criteria={"billing": "", "sales": "", "retail": ""}),
}
result = engine.ask("Help! My payouts have been failing for 3 days!", questions)
print(result.answers["team"].choice, result.answers["team"].confidence)
print(result.answers["urgent"].noul)

Set device to match your hardware: "cuda" for an NVIDIA GPU (it is the default, so it fails on machines without one), "mps" for an Apple Silicon Mac, or "cpu". Choice answers have .choice, .confidence and .probabilities; yes/no answers have .noul; score answers have .score. The package is version 0.1.0, so expect names to settle in later releases.

Load the model once and reuse it for every decision; loading takes seconds, deciding takes milliseconds. The pattern Ethan used for Jake's form is the useful one in practice: if the confidence for the chosen option is above a threshold you pick (say 0.8), act on it automatically; below it, send the item to a person. Start strict, and loosen the threshold as you see how it does on your real data.

We tested it: a 2016 laptop, CPU only

To check every command in this guide, we installed Strands Decider on an ordinary older laptop: Windows 11, 16 GB of RAM with Python 3.14 from 2016.

StepResult
pip install torch (CPU build) and strands-deciderWorked first time, no compiler needed
First strands-decider ask (includes download)Worked; 15 minutes, almost all of it downloading 4.4 GB of model files
Answers to the model card examplebilling (0.844), urgency 0.829, "frustrated" (0.572), matching the published example
Loading the model in PythonAbout 19 seconds, once
Each decision (two questions)4.9 to 5.3 seconds, median 5.1 seconds
strands-decider serve and the curl request aboveReady in about 15 seconds on 127.0.0.1; correct JSON answer in 2.6 seconds
README Python exampleFails in package 0.1.0; use the corrected code in the Python section

So the honest summary: on a recent GPU it is a real-time model, about a tenth of a second per decision; on an old CPU it is about forty times slower, which still suits batch jobs and low volumes. The confidence numbers behaved well in our test. Eight sample messages produced high confidence (0.77 to 0.86) only for the three clear billing requests, all routed correctly, and low confidence (0.17 to 0.46) for the ambiguous ones, including a thank-you message that fit no team. With the 0.8 threshold suggested above, only correct answers would have been acted on automatically.

Two warnings you will see, and what they mean

  • "cache-system uses symlinks by default … your machine does not support them" (Windows only). Hugging Face's download cache prefers symbolic links, which Windows allows only in Developer Mode or for administrators. The cache still works, just using a little more disk. Turn on Developer Mode in Settings > System > For developers to remove it, or set the environment variable HF_HUB_DISABLE_SYMLINKS_WARNING=1 to hide it.
  • "causal_conv1d_fn is falling back to its reference PyTorch implementation" and "flash-linear-attention is not installed". The Qwen3.5 base model uses some newer layer types that have optional fast kernels, mostly for NVIDIA GPUs on Linux. Without them, the answers are exactly the same, just slower, which is part of why CPU decisions take seconds. On a CPU-only Windows PC there is nothing to install; ignore them.

Two practical tips from the test: the first run downloads about 4.4 GB, so if your system drive is nearly full, point the download elsewhere first by setting the HF_HOME environment variable to a folder on another drive. And on Windows, the PyTorch package that pip installs by default is the CPU version, which is what you want unless you have a supported NVIDIA card and driver.

With Strands Agents and other AI agents

Strands Decider was built for agents. The Strands Agents SDK, AWS's open-source framework for building AI agents, lets you run code just before an agent calls a tool. The project's example uses that moment to ask the decider two yes/no questions before a weather tool runs, and if the user never said which city, the agent asks instead of guessing.

That is the general idea: a fast, cheap gate in front of slow or risky actions. Is this request in scope? Does the user's message contain an email address the agent should not send to? Is this the kind of question the support agent should hand to a human? A full language model can answer those too, but slower, at higher cost, and less predictably. If you are building agents on AWS, our explainer on Bedrock AgentCore shows where such checks fit in a production setup.

Is it on Amazon Bedrock? Do you need an AWS account?

  • No AWS account is needed. The model and package download from Hugging Face and PyPI and run entirely on your machine, with no internet once the weights are downloaded.
  • Not on Bedrock as of early October 2026. It is an open model you host yourself. You can of course run it on an EC2 instance or in a container on AWS, like any Python service.
  • Costs: nothing for the model. Running it on your own computer costs electricity; on a cloud server, the server's hourly price.

If you want hosted AI models on AWS instead, our plain-English guide to Amazon Bedrock covers what is available there.

What hardware do you need?

HardwareWhat to expect
NVIDIA RTX 3090 or similarAbout 115 ms median per decision, 299 ms at the slow end
Apple M3 ProAbout 153 ms median for short inputs, 234 ms across all tasks
Smaller NVIDIA cards (8 GB and up)Should run comfortably; a 2B model needs a few GB of VRAM
CPU only (our test: Intel Core i7-6700HQ, 2016)Works, but about 5 seconds per decision with two questions, and 2.6 seconds for one. Fine for a few hundred decisions a day, not for real-time use

Compared with chat models, this is light: there is no long answer to generate, and the input is capped at 4,096 tokens. Our guide to hardware tiers for local AI is useful if you plan to run larger models alongside it.

Train your own version

Unusually, AWS published the full training recipe, data sources and evaluation harness. If the general model does not fit your categories well, you can fine-tune it on your own labeled examples. The recipe runs on a single 24 GB GPU such as an RTX 3090 in about 11 hours, or on eight H100 GPUs (for example an AWS p5.48xlarge instance) in about 70 minutes. The model card's advice is worth following: yes/no and score questions transfer less well to new domains than choices do, so if those matter to you, train on your own rubric.

Limitations to know before you rely on it

  • It cannot write. No summaries, explanations, code or replies. It is a gate in front of other tools, not a replacement for them.
  • Short inputs only. 4,096 tokens of context, and long, multi-step documents perform noticeably worse.
  • Wording matters. Rephrasing a question, or swapping how options are labeled, can change answers. Fix your wording and test it.
  • Calibration was fitted on short classification tasks, so confidence numbers are most trustworthy for that kind of question.
  • Inherited bias. It was trained on public datasets and inherits their domains and label noise. Check it on your own data before automating anything that affects people.
  • Distracting text hurts. Irrelevant context in the state lowers accuracy; send only what the decision needs.

Community versions: ONNX, WebGPU and GGUF

Within days of release, the community published conversions: an ONNX version (useful for running it from other languages and runtimes), a WebGPU build that runs the model inside a web browser, and a GGUF repack. Treat them as experiments. The decision depends on the pointer head, and repacks of other decision models have sometimes carried the base model without that head, leaving a model that cannot actually answer typed questions. The official Python package is the supported route; if you try a conversion, compare a few answers against the official package first.

Who should use Strands Decider?

  • Developers building agents who want a fast, cheap guardrail before tool calls.
  • Small businesses routing messages, tickets, form entries or emails, on a PC they already own.
  • Teams with privacy rules who cannot send customer text to a hosted API.
  • Anyone comparing decision models, since it speaks the same API as Jev, Clef in Ollama and llama.cpp's decision server.

Who should look elsewhere: anyone who needs written answers (use a chat model), questions about images (Clef reads images), or long documents (larger models handle them better).

AWS Strands Decider 2B: frequently asked questions

What is AWS Strands Decider 2B?

An open-source decision model from AWS Strands Labs that answers yes/no, choice and score questions about text with calibrated confidence, without generating any text.

Is Strands Decider free?

Yes. It is released under Apache 2.0, free for commercial use, and runs on your own hardware.

How do I install Strands Decider?

Create a Python virtual environment and run pip install strands-decider. The model downloads from Hugging Face on first use.

Do I need an AWS account to use Strands Decider?

No. It runs locally with no AWS account, and offline once the model is downloaded.

Is Strands Decider available on Amazon Bedrock?

Not as of early October 2026. You host it yourself, on your own computer or a server.

How fast is Strands Decider 2B?

About 115 milliseconds per decision on an RTX 3090 and about 150 milliseconds on an M3 Pro Mac for short inputs.

How accurate is Strands Decider?

It scores 0.723 on the public JevBench set, the best of the true 2B models and third of 33 overall, and 0.641 on held-out short tasks.

What is the difference between Strands Decider and Jev?

Jev is TypeSafe's closed, hosted decision model. Strands Decider is open, small and self-hosted, and uses the same /v1/systemone request format.

How does Strands Decider compare with Cloudflare Clef?

Clef is larger (9B and 27B) and reads images. Strands Decider is smaller, faster and text-only. Both are Apache 2.0.

Can Strands Decider run on a CPU?

Yes. In our test on a 2016 laptop, each decision took about 5 seconds on the CPU, versus about 0.1 seconds on a modern NVIDIA GPU.

What question types does Strands Decider support?

Three: yes/no (called noul), choice between options, and score on an ordered scale.

What is /v1/systemone?

The request format for decision models, introduced with Jev and now used by Ollama's decision models, llama.cpp and Strands Decider's server.

How do I run the Strands Decider server?

Run strands-decider serve with the model name and a port, then send JSON requests to /v1/systemone on that port.

Can Strands Decider write text or summaries?

No. Its text head was removed; it can only choose among the options you provide.

What is the context length of Strands Decider?

4,096 tokens. Keep inputs short and relevant for the best accuracy.

Can I fine-tune Strands Decider on my own data?

Yes. AWS published the training recipe; it takes about 11 hours on an RTX 3090 or about 70 minutes on eight H100s.

How do I install Strands Decider on Kali Linux?

Create a venv with python3 -m venv, activate it, and pip install strands-decider inside it, or use pipx for the command-line tool.

Is Strands Decider in Ollama?

Not officially as of early October 2026. Use the official Python package and server.

Is the Strands Decider server secure?

It has no authentication. Keep it on 127.0.0.1, or put it behind your own access control.

What is Strands Decider built on?

Qwen3.5-2B-Base, fine-tuned with LoRA, with a small pointer head in place of the language-modeling head.

What does calibrated confidence mean?

That the confidence number matches reality: answers given with 0.8 confidence are right about 80% of the time, so you can set thresholds.

Why does "from strands_decider import DeciderModel" fail?

That name does not exist in the released package 0.1.0. Use load_engine from strands_decider.infer with the question classes from strands_decider.schema, as shown in the Python section.

What does the "flash-linear-attention is not installed" warning mean?

Optional fast kernels for the base model's newer layers are missing, so slower but identical code runs instead. On a CPU-only PC, ignore it.

Why does Hugging Face warn about symlinks on Windows?

Its cache prefers symbolic links, which need Developer Mode on Windows. The cache still works; turn on Developer Mode or set HF_HUB_DISABLE_SYMLINKS_WARNING=1.

How big is the Strands Decider download?

About 4.4 GB on first use, mostly the Qwen3.5-2B base model. Set HF_HOME to store it on another drive.

What is Strands Agents?

AWS's open-source SDK for building AI agents. Strands Decider can gate an agent's tool calls with quick yes/no checks.

Strands Decider 2B is a small model with a narrow, valuable job: quick, confident decisions on text, with no writing, no hosted API and no bill. It is open, fast enough for real-time use on ordinary hardware, speaks the same API as the other decision models, and comes with everything needed to train your own. It will not replace a chat model, and it is not meant to. Jake's booking form now routes itself in a tenth of a second, and the messages it is unsure about still reach a person, which is exactly how it should be.

 If you keep one line from this page

pip install strands-decider in a venv, then strands-decider serve: a free, local /v1/systemone decision server in two commands.

Act on confident answers; send the uncertain ones to a person.

Revision note. Written October 4, 2026, days after AWS released Strands Decider 2B, and tested the same day on a 2016 laptop. May your quick decisions be the right ones.

Related