Open-Source Jev AI Models: JEV-9B, JEV-27B-VL and GEV-26B-Decide Explained (Download, Hardware, AWS, the GGUF Trap)
The open-source Jev models are free, downloadable copies of TypeSafe's hosted Jev decision model, published on Hugging Face by an independent group called AutoTrust AI: JEV-9B, JEV-27B, JEV-27B-VL and GEV-26B-Decide, all under the Apache 2.0 license, so you can use them commercially. They answer the same typed questions Jev does (yes/no, pick one option, rate 0 to 5) with a probability for every answer, and they can also chat and reason like a normal model. The shock is in the files people actually download: the popular GGUF versions of JEV-9B, JEV-27B and GEV-26B-Decide contain no decision part at all. They are plain Qwen and Gemma models with a Jev name on the label, because the part that makes decisions lives in separate small files that the GGUF conversion leaves behind. Below: what each model is, which one fits your hardware, how to spot a hollow download, how to run the real thing on Windows, Kali or AWS, and when the hosted Jev at $0.042 per million tokens is simply the better deal.
Jake had heard about Jev from Priya, Ethan's friend, who now sorts her shop's email with decision models. He wanted the same for the repair shop: read each customer message, answer "is the screen cracked, yes or no," and route it. On Saturday he found "GEV-26B-Decide GGUF" on Hugging Face, saw 400,000 downloads on the original repository, downloaded the 17 GB Q4_K_M file into LM Studio, and typed his question. He got three friendly paragraphs about screen damage, a suggestion to visit a repair shop, and no probability anywhere. He sent Ethan a screenshot: "Is it broken or am I?" Ethan looked at the file list for ten seconds and wrote back: "Neither. You downloaded Gemma. The decision part isn't in that file."
The basics: Jev, decision models, and "System 1"
If you have used an AI chatbot, you have used a model that writes: it produces its answer one word at a time. A decision model is different. You give it a situation, called the state, and a typed question, and instead of writing, it returns a probability for every possible answer in a single pass. Software loves this, because there is nothing to parse: the email is 97% "refund request," the screenshot is 91% "login page," and your code can act on the number.
Jev is the model that made the idea famous. TypeSafe AI launched it on September 15, 2026 as the first of what it calls "System One" models, named after the fast, intuitive mode of thinking in psychology. Jev answers three kinds of question:
noul, a yes/no question, returning probabilities forfalseandtrue. ("Is the customer asking for a refund?")choice, a pick from your own list of options. ("Which team should handle this: billing, shipping or tech support?")score, a rating from 0 to 5. ("How urgent is this message?")
Jev itself is hosted and closed: you call TypeSafe's API, and the weights are not public. It costs $0.042 per million input tokens, with output free, which is why so many people use it for sorting and routing at scale.
If you searched for "Jev AI GitHub," here is what exists. TypeSafe's GitHub organization, typesafe-ai, publishes its client libraries under the MIT license, such as typesafe-sdk-python, plus system-one-adapter-python, a drop-in replacement client that answers the same questions using ordinary LLM APIs. None of those repositories contain Jev's weights. The open models people mean when they say "open-source Jev" live on Hugging Face, and they are the subject of this page.
"So it's a filing clerk," Jake said. "It doesn't write the letter, it decides which drawer the letter goes in."
"And it tells you how sure it is," Ethan said. "That last part is the whole point."
System 1 and System 2 in one model
The JEV models add a twist. Each one is built on a normal chat model, and it keeps that ability. So one download gives you two ways of thinking, and every request chooses one:
- System 1 is the fast decision mode: one pass, a calibrated probability distribution, well under a second.
- System 2 is ordinary text generation with step-by-step reasoning, exactly as the base model does it.
That pairing enables a pattern AutoTrust calls confidence-gated escalation: let System 1 answer when it is sure, and hand the unsure cases to System 2 to think through. We show the code for it further down.
The open-source Jev family: four models, one recipe
All four come from AutoTrust AI, which is not affiliated with TypeSafe. AutoTrust trained them as "students" on a public, Apache-2.0 dataset of Jev's own output distributions, so their System 1 imitates Jev closely. Here is the family as of October 6, 2026:
| Model | Built on | Download | Sees images? | Best at | Released |
|---|---|---|---|---|---|
| JEV-9B | Qwen3.5-9B | about 18 GB | Yes, since October 3 | Speed; the smallest and fastest | September 23 |
| JEV-27B | Qwen3.8-27B | about 54 GB | No (text) | Closest to Jev; long option lists | September 25 |
| JEV-27B-VL | Qwen3.8-27B with vision | about 54 GB | Yes | Screenshots, photos, 256K-token inputs | September 30 |
| GEV-26B-Decide | Gemma 4 26B-A4B (about 4B active) | about 49.5 GB, or 18 GB as NVFP4 | Yes | Adaptive thinking on hard questions | September 29 (renamed October 2) |
They have been popular. At the time of writing, JEV-27B-VL had about 1.28 million downloads, GEV-26B-Decide about 446,000, JEV-9B about 305,000 and JEV-27B about 134,000. GEV first appeared as "JEV-Gemma4-26B-A4B" and was renamed on October 2, so older links may use that name.
The "Blocks of Experts" recipe, and why it matters
AutoTrust did not retrain the base models. Its recipe, which it calls Blocks of Experts, keeps the pretrained model frozen and adds a small, detachable block trained only for decisions. In JEV-9B, that decision block is 40.2 million trained parameters, about 0.5% of the backbone, trained in roughly three hours on a single B200 GPU. In JEV-27B, the decision block is about 109 million parameters, roughly 0.4% of the model, and it trained in about 9.2 hours on one B200; its adapter file is 416 MB.
Keeping the blocks separate has a measurable payoff. JEV-9B scores 70.7% on the HumanEval coding test with the decision block attached, identical to the base Qwen3.5-9B, and all 164 of its answers are byte-identical to the base model's. Folding the same block into the backbone instead would have cost nine points (61.6%), by AutoTrust's own test.
It also explains the trap in the next section. In JEV-9B, JEV-27B and GEV, the main weight files are the original base model, bit for bit. Everything that makes the model a Jev lives in a few small extra files: the adapter, a decision "head," and a calibration file.
How close to Jev are they?
AutoTrust measures closeness with KL divergence, a standard way of comparing two probability distributions, where 0 means identical. On 25,376 held-out questions from 53 domains, labeled with Jev 1.13's own outputs:
- JEV-9B scores about 0.019, agrees with Jev's top choice 90.2% of the time, ranks yes/no answers with an AUROC of 0.994, and keeps 90% of Jev's accuracy on an independent 16-option test.
- JEV-27B scores about 0.017, keeps 96% of Jev's accuracy at 16 options, and more than halves the error on task types it never trained on (0.234 down to 0.104).
Those are AutoTrust's own measurements, but they are unusually specific, and the dataset is public, so anyone can check them.
The GGUF trap: most JEV downloads have no decision part
GGUF is the file format used by llama.cpp, LM Studio and Ollama, and for most models a GGUF is the easiest way to run locally. For the JEV family, it is mostly a trap. Here is what is in the popular community GGUF repositories:
| GGUF repository | What is inside | Makes Jev decisions? |
|---|---|---|
| mradermacher/GEV-26B-Decide-GGUF (Q4_K_M 16.9 GB, Q8_0 27.0 GB) | Only .gguf weights and a vision file: plain Gemma 4 26B-A4B | No |
| prithivMLmods/JEV-27B-GGUF (Q4_K_M 16.5 GB) | Plain Qwen3.8-27B | No |
| prithivMLmods/JEV-9B-GGUF and mradermacher/JEV-9B-GGUF (Q4_K_M 5.63 GB) | Only .gguf weight files; no adapter, head or calibration: plain Qwen3.5-9B | No |
| Atlas3D/JEV-27B-VL-GGUF (Q4_K_M 16.55 GB, Q8_0 28.6 GB) | Weights with the decision readout, plus a llama.cpp patch | Yes, with a patched llama.cpp |
Why does this happen? A GGUF converter takes the main weight files and compresses them. In this family, those main files are the untouched base model. The adapter, the 24-slot decision head and the calibration temperatures sit beside them as separate small files, and a standard conversion simply does not carry them. The result loads perfectly, chats perfectly, and answers your yes/no question with three paragraphs, exactly as Jake saw.
"So the label says Jev," Jake said, "and the box has a very nice Gemma in it."
"A very nice Gemma," Ethan agreed. "Just not the one you ordered."
How to tell a real JEV download from a hollow one
- Look at the file list, not the name. A real JEV repository has
head.safetensors, anadapter/oradapter_vllm/folder, andcalibration.json. If a repository holds only.gguffiles (and maybe anmmprojvision file), it cannot make Jev decisions. - Check for a patch. The one working GGUF, Atlas3D's JEV-27B-VL, ships a
llama.cpp-jev.patch, because the "jev" decision type is not in upstream llama.cpp. - Ask a yes/no question. A real System 1 returns numbers for
falseandtrue. A hollow model writes sentences. - Don't import it into Ollama and hope. Ollama's library has no JEV models, and a hand-imported hollow GGUF has no head, so there is nothing for Ollama's
/v1/systemoneendpoint to score.
Why 4-bit compression hurts decisions more than chat
Even the real GGUF loses something at 4-bit. Atlas3D measured it: the Q8_0 file (28.6 GB, about 33.6 GB of graphics memory in use) agrees with the full-precision model on 99.0% of decisions (306 of 309). The Q4_K_M file (16.55 GB, about 22.7 GB in use) agrees on 92.9%, so about one decision in fourteen comes out differently.
For chat, that would barely be noticeable; a slightly different word choice is still a fine sentence. For decisions, it is the answer itself that changes. If you use a 4-bit build, test it on a sample of your own messages against the full model before you trust its thresholds.
Is the Jev model free? The license, in plain words
Yes: all four JEV models are free to download and use commercially, under Apache 2.0. That follows from their parts. The base models are Apache 2.0 (Qwen3.5-9B and Qwen3.8-27B from Qwen, and Gemma 4, which Google released under Apache 2.0), and the training dataset, SargeDev/jev-distill-corpus-v3, is Apache 2.0 too, with one of its streams additionally released as CC0.
That is a real difference from the model most people find first. OpenJev, the 27B model that llama.cpp supports through its /v1/systemone endpoint, is licensed CC BY-NC 4.0, which rules out business use. If you want a Jev-style model for a shop, a product or a client, the JEV family and the small Apache models in our OpenJev and llama.cpp guide are the legal routes.
Two honest footnotes. Free weights do not mean free running: you need a capable GPU or cloud time, which we price below. And the System 1 idea and the noul, choice and score primitives come from TypeSafe's Jev; AutoTrust's models are independent students that share no weights or code with it, and are not endorsed by TypeSafe.
What hardware you need for each JEV model
The decision part is tiny. The base model is not, and it has to be loaded in full. Here is what each one needs, from the model cards and the file sizes:
| Model | GPU memory | Example cards |
|---|---|---|
| JEV-9B (full precision, 18 GB) | 24 GB is the practical floor for short decision prompts; 32 to 48 GB is comfortable | RTX 4090 or 3090 (24 GB), RTX 5090 (32 GB), L40S (48 GB) |
| GEV-26B-Decide NVFP4 (17.1 GiB) | 24 GB with a short context; 32 to 48 GB for long contexts and many users | RTX 5090, RTX PRO 6000, DGX Spark (native FP4); H100 or A100 (fallback mode) |
| JEV-27B, JEV-27B-VL, GEV-26B-Decide (full precision) | One GPU with 80 GB or more | H100 80 GB, RTX PRO 6000 (96 GB), H200, B200 |
| JEV-27B-VL Q8_0 GGUF (patched llama.cpp) | About 33.6 GB in use | 48 GB cards; 99.0% agreement with full precision |
| JEV-27B-VL Q4_K_M GGUF (patched llama.cpp) | About 22.7 GB in use | 24 GB cards; 92.9% agreement |
Speed on the cards' own test hardware (one B200 data-center GPU): JEV-9B answers a single decision in a median of about 90 milliseconds and sustains about 340 decisions per second on an independent benchmark; JEV-27B takes about 137 milliseconds; GEV's System 1 about 45 milliseconds. Atlas3D measured about 80 milliseconds per decision for its Q8_0 GGUF on an RTX PRO 6000. A consumer card will be slower than a B200, but even several times slower is still well under a second.
Windows, Mac and Kali Linux
- Windows 11: the official serving route is vLLM, which runs on Linux. Use WSL2 with Ubuntu, and install the NVIDIA driver on the Windows side; WSL2 passes the GPU through. Our WSL guide covers the install.
- Kali Linux and Ubuntu: vLLM runs natively with an NVIDIA GPU. On Kali, install it inside a Python virtual environment; Kali blocks system-wide pip installs, as explained in our guide to the externally-managed-environment error.
- Mac: there is no published Mac route for the decision part today. vLLM's JEV setup needs NVIDIA's CUDA, and the hollow GGUFs that do run on a Mac do not make decisions. A Mac owner who wants local decisions should use the small models in Ollama or llama.cpp covered in the alternatives section.
- No GPU at all: the full JEV models are not practical on a CPU. AWS's Strands Decider 2B and the 0.8B Tev1 model run on ordinary processors.
How to run JEV-9B, step by step
JEV-9B is the easiest member to run, and its newer vision server gives you the simplest interface: a /v1/decide endpoint that takes plain JSON and returns probabilities, no client-side math needed. It also answers text-only questions, matching the text model to within 0.011 on 300 test decisions. These steps work on Ubuntu, Kali, or Ubuntu inside WSL2 on Windows.
- Check the GPU. Run
nvidia-smi. You should see your card and at least 24 GB of memory. If the command is missing inside WSL2, update the NVIDIA driver on Windows first. - Create a virtual environment.
python3 -m venv ~/jev-env source ~/jev-env/bin/activate pip install -U huggingface_hub requests
- Install a recent vLLM development build. The JEV cards were tested with a vLLM development build from September 2026, because they need Qwen3.5 support, LoRA on the output layer, and a couple of newer options. At the time of writing, vLLM's nightly wheels install with:
pip install -U vllm --pre --extra-index-url https://wheels.vllm.ai/nightly
- Download the vision server files.
hf download autotrust/JEV-9B --include "vl/*" --local-dir JEV-9B
- Start the server.
bash JEV-9B/vl/serve.sh
The script downloads Qwen/Qwen3.5-9B with its vision encoder and serves both systems on port 8000. Be patient: start-up takes three to eight minutes while vLLM prepares the GPU, and it looks frozen while it does. Keep its--max-num-seqs 8setting. - Ask your first question, shown in the next section.
If you only need text and want the leaner setup, the card's main route downloads the full repository (about 18 GB) and runs vllm serve with the decision adapter attached as a LoRA module named jev-decision, plus --logprobs-mode processed_logprobs, which is required for correct probabilities. That route asks you to apply the head's bias and temperature in your own code; the card includes a ready-made Python function for it.
Your first decision request, and confidence gating
With the server running, a yes/no question is one HTTP call. This works from any language; here it is with curl:
curl localhost:8000/v1/decide -H 'Content-Type: application/json' -d '{
"kind": "noul",
"state": "Customer: dropped my phone yesterday, now there are lines across the display and a crack in the corner.",
"question": "Is the screen cracked?"}'
And a choice question, the kind Jake wanted for routing:
curl localhost:8000/v1/decide -H 'Content-Type: application/json' -d '{
"kind": "choice",
"state": "Customer: my card was charged twice for one screen repair.",
"question": "Which desk should handle this?",
"options": ["billing", "repairs", "sales"]}'
The response lists your options with a probability for each, in the same order, plus the winning choice. AutoTrust's own example for a double-charged coffee returned 0.9969 for "billing," 0.0000 for "shipping" and 0.0031 for "tech support." For noul, the options are always ["false", "true"]; for score, they are 0 through 5. JEV-9B accepts 2 to 16 options per choice question, and JEV-27B, JEV-27B-VL and GEV accept up to 256.
Two input rules save surprises. Keep states under about 1,024 tokens on JEV-9B, or set a higher limit, because longer inputs are trimmed by default (keeping the first 60% and the last 40%). And for images, send a list that mixes text and image parts, such as ["Photo:", {"image": "data:image/png;base64,..."}].
Confidence gating: let System 1 handle the sure cases
The calibration is what makes the numbers useful: AutoTrust reports an expected calibration error of 0.0007 for JEV-9B, meaning a stated 80% really is right about 80% of the time on its test set. That lets you build a simple rule:
- Ask System 1.
- If the top answer is above a threshold you picked, such as 0.90, act on it automatically.
- Otherwise, send the same question to System 2 (the normal chat endpoint, with thinking on) or to a person.
The JEV-9B card includes a ready-made solve() function that does exactly this against the same server. Two cautions from AutoTrust itself: the threshold must come from your own validation data, not from a blog post, and System 2 was not trained to agree with System 1, so they will sometimes disagree. When they do, that disagreement is itself a useful flag for human review.
That is the shape of Jake's plan: System 1 for the routine "is my phone ready" and "how much for a screen" messages, and a person for anything it is unsure about. Before he picks the threshold, he will run a few hundred of his own past messages through it and see where its confidence and his judgment part ways.
How to run the JEV models on AWS
No GPU at home? AWS rents single-GPU machines by the hour, and every JEV model fits on one GPU. The models are not on Amazon Bedrock, so this means an EC2 instance (a rented virtual server) running the same vLLM setup as above. These are the instances that fit, with AWS's on-demand Linux prices in US East (N. Virginia) on October 6, 2026:
| Instance | GPU | GPU memory | Runs | Per hour | Per month, 24/7 |
|---|---|---|---|---|---|
| g6.xlarge | 1× NVIDIA L4 | 24 GB | JEV-9B with short prompts (tight) | $0.80 | about $590 |
| g5.xlarge | 1× NVIDIA A10G | 24 GB | Same as g6.xlarge, older and pricier | $1.01 | about $730 |
| g6e.xlarge | 1× NVIDIA L40S | 48 GB | JEV-9B comfortably; JEV-27B-VL Q8_0 GGUF | $1.86 | about $1,360 |
| g7e.2xlarge | 1× RTX PRO 6000 Blackwell | 96 GB | All four at full precision; GEV NVFP4 natively | $3.36 | about $2,455 |
| p5.4xlarge | 1× NVIDIA H100 | 80 GB | The 80 GB models; faster memory | $6.88 | about $5,020 |
The standout is g7e.2xlarge: at about $3.36 an hour, it is the cheapest single GPU on AWS with 80 GB or more, so it runs JEV-27B, JEV-27B-VL and the full GEV-26B-Decide, and its Blackwell GPU runs GEV's NVFP4 build in native FP4. For JEV-9B alone, g6.xlarge at about $0.80 an hour is the budget pick, with the caveat that 24 GB is tight. AutoTrust tested its models on a B200, so treat these AWS pairings as the memory math, not as published benchmarks, and confirm on a short trial before you commit.
On an 80 GB GPU such as the H100 in p5.4xlarge, the JEV-27B cards advise starting with --max-model-len 131072 rather than the full 256K, because the context eats into the spare memory. The xlarge sizes also come with only 16 to 32 GB of ordinary system memory; if model loading stalls, the next size up gives more.
Step by step on EC2
- Check your quota. In the Service Quotas console, look up Running On-Demand G and VT instances (for g6, g5, g6e and g7e) or Running On-Demand P instances (for p5). Quotas are counted in vCPUs: g6.xlarge needs 4 and g7e.2xlarge needs 8. New accounts often start at zero for GPU families, so request the increase a day ahead. Our guide to the vCPU quota exceeded error shows how.
- Launch the instance with an AWS Deep Learning AMI for Ubuntu, which includes the NVIDIA driver and CUDA. Give it a gp3 EBS volume of about 150 GB for JEV-9B, or 250 GB for the 27B models, so the weights, the base model and the Python packages all fit.
- Install and serve exactly as in the JEV-9B steps: a virtual environment, the vLLM nightly build,
hf download, then the model'sserve.sh. For JEV-27B-VL and GEV, download the full repository and run itsserve.sh; GEV also needs the vLLM patch bundled in its repository. - Keep port 8000 private. Do not open it in the security group. Reach it through an SSH tunnel or Systems Manager port forwarding:
aws ssm start-session --target <your-instance-id> \ --document-name AWS-StartPortForwardingSession \ --parameters '{"portNumber":["8000"],"localPortNumber":["8000"]}'Then your scripts callhttp://localhost:8000/v1/decideas if the server were on your desk. - Stop it when idle and set a budget alert. A stopped instance costs only its storage. Our three-layer billing alert setup catches a forgotten GPU before the bill does.
Amazon SageMaker AI can host the same models as a managed endpoint if your team already works there; you pay more per hour for the managed serving. Our Bedrock vs SageMaker AI guide explains the trade.
Self-host on AWS or use hosted Jev? The honest math
Hosted Jev costs $0.042 per million input tokens with free output. A g6.xlarge running all month costs about $588, which buys about 14 billion input tokens of hosted Jev. At 500 tokens per decision, that is roughly 920,000 decisions a day before self-hosting even breaks even on the cheapest instance. For g7e.2xlarge, the break-even is closer to 3.8 million decisions a day.
So for most small businesses, hosted Jev is cheaper. Self-hosting a JEV model makes sense for other reasons:
- Privacy: the messages, screenshots or documents never leave your machine or your AWS account.
- Control: no rate limits, no model changes you did not choose, and the option to fine-tune the decision block on your own labels.
- Huge or bursty volume that keeps a GPU busy around the clock.
- System 2 in the same box: the same server also writes and reasons, which hosted Jev does not do.
"So I'm not saving money," Jake said.
"Not at your volume," Ethan said. "You'd do it so customer messages stay in the shop. That's a fine reason. Just know it's the reason."
Which JEV model should you pick?
- You have a 24 GB NVIDIA card and want text and image decisions: JEV-9B. It is the fastest of the family and the simplest to run.
- You need long option lists, unfamiliar task types or code-rule checks: JEV-27B. AutoTrust itself recommends it over JEV-9B for these, and it keeps 96% of Jev's accuracy at 16 options.
- You judge screenshots, photos or very long documents: JEV-27B-VL, which reads images and prompts up to 256K tokens. On a 48 GB card, Atlas3D's Q8_0 GGUF with its patched llama.cpp is the alternative.
- Your questions are hard, like puzzles, math or game states: GEV-26B-Decide with adaptive thinking, below. On a Blackwell card with 24 to 32 GB, its NVFP4 build fits.
- You have a laptop, a Mac or no NVIDIA GPU: none of these. Use the smaller open models in the alternatives section.
- You make fewer than a few hundred thousand decisions a day and privacy is not a hard rule: hosted Jev.
GEV-26B-Decide and adaptive thinking
GEV is the family's most interesting experiment. Built on Google's Gemma 4 26B-A4B, a mixture-of-experts model that uses about 4 billion parameters per token, it adds adaptive thinking to System 1. Send "thinking": "auto" with a decision, and GEV answers instantly when its confidence clears a threshold (0.8 by default), but stops to reason first when it does not. Thinking costs time, roughly one second per 250 thinking tokens, so it is spent only where it pays.
Where it pays is striking. On AutoTrust's game and puzzle tests, adaptive thinking took Minesweeper from 21.0 to 86.0, Connect Four from 54.0 to 99.5, Wordle from 52.0 to 100, and Sudoku from 76.0 to 99.5. And where it does not pay, AutoTrust says so: on classification tasks the model already knows, thinking made things slightly worse (BANKING77 dropped from 88.0 to 85.0, CommonsenseQA from 86.7 to 85.0). On the very hard Humanity's Last Exam questions, System 1 alone answered just 1.6% correctly, a reminder that a fast guess is not a substitute for real reasoning.
AutoTrust also reports a Decision Index score of 62.48 for GEV with adaptive thinking, against 57.91 for Jev 1.13 on the public board. That number comes from AutoTrust's own scoring run, not a leaderboard entry, so read it as promising rather than proven. On speed, GEV's System 1 takes about 45 milliseconds on a B200 and sustains about 257 decisions per second with 64 clients.
The NVFP4 build released on October 5 shrinks the download to 18 GB (17.1 GiB of weights). It runs natively on Blackwell GPUs such as the B200 (tested), RTX PRO 6000, RTX 5090 and DGX Spark (not yet tested), and in a slower fallback mode on H100, H200 and A100. On the GPQA Diamond science test, its System 1 scored 44.9 against 43.9 for full precision, which is within noise.
The robot arm and computer-use demos, honestly
The JEV models went viral partly on videos of a robot arm and a web browser driven by decisions. Here is what the numbers behind them say, across the family, all in simulation or a headless browser:
| Test | JEV-27B-VL | JEV-9B | GEV-26B-Decide |
|---|---|---|---|
| Robot arm pick and place (20 simulated scenes) | 75% | 50% | 40% |
| Computer use, numbered boxes plus element text (60 tasks) | 95% | 95% | 95% |
| Computer use, numbered boxes only | 10% | 37% | 15% |
| Time per robot-arm decision | 239 ms | 163 ms | 61 ms |
Two lessons stand out. The models are strong judges inside a loop that asks simple questions ("is the cube left or right of the gripper?"), but asked to pick one of eight motor commands directly, JEV-27B-VL went 0 for 10. And for computer use, they need the text of each button or link, the way an accessibility tree provides it; with numbered boxes alone, success collapses.
JEV-27B-VL's other results are quietly more useful for businesses: 89.5% accuracy picking among all 150 intents of the CLINC150 benchmark with no training (93.8% when each intent has a description), 20 of 20 correct on decisions that hinge on one sentence hidden in up to 250K tokens of text, and 73.2% on the Plan-RewardBench agent-judging test, ahead of GPT-5's 68.5% in that paper's table. Recommending short videos from cover images alone, it matched a collaborative-filtering system trained on 59,045 users' viewing logs (AUC 0.727 against 0.728).
Limitations AutoTrust admits, and a few more
The JEV model cards are refreshingly frank about their weak spots. The ones that matter most in practice:
- They copy Jev's mistakes along with its strengths. A student trained to match a teacher's probabilities inherits the teacher's errors. Published evaluations of Jev show it is unreliable on multi-step reasoning, arithmetic, dates, counting and deliberately tricky inputs, and the JEV models share all of that. Do not ask a decision model to add up an invoice.
- JEV-9B adds blind spots of its own. On 16-option questions, 11.5% of its answers change when only the order of the options changes (Jev: 7.0%). Shuffle your options in testing to see how stable your own questions are.
- System 1 cannot explain itself, and System 2 does not know what System 1 decided. Expect occasional disagreement when you escalate.
- English first. The training data is English; other languages work only as well as the base model happens to handle them, untested.
- One part of the training data teaches nothing. The dataset's memory-relevance rows are all exactly 50/50 placeholders, so the models answer about 0.5 on those questions by design. Do not use them to score memory relevance.
- Speed comparisons with hosted Jev are not like for like. AutoTrust's timings exclude network time; the hosted figures, measured by third parties, include it.
- Not for high-stakes decisions. AutoTrust's own advice: act automatically only above a threshold validated on your data, and route the rest to a stronger model or a person.
- The software is young. The serving setup depends on a vLLM development build and, for GEV, a bundled patch. Expect the steps to change as vLLM catches up.
Troubleshooting JEV models
| Problem | Likely cause | Fix |
|---|---|---|
| The model writes paragraphs instead of probabilities | You loaded a hollow GGUF, or sent a chat request instead of a decision | Use the official repository and /v1/decide, or the patched Atlas3D GGUF |
| vLLM rejects the model or a flag | Your vLLM release is too old for Qwen3.5 or Gemma 4 with LoRA on the output layer | Install the nightly build; for GEV, apply the bundled patch |
| The server seems frozen at start | CUDA-graph capture with LoRA takes three to eight minutes | Wait; watch nvidia-smi for activity |
| CUDA out of memory | Context length too large for the card | Lower --max-model-len or MAX_MODEL_LEN; use a bigger GPU |
| JEV-27B-VL probabilities look wrong under load | More than 8 sequences per batch breaks its LoRA path | Keep --max-num-seqs 8; extra requests simply queue |
| Probabilities do not add up or look truncated when calling completions directly | Default top_k=20 and top_p=0.95 cut off options | Pass top_k: 0 and top_p: 1.0, or use /v1/decide |
| Image decisions are slow | Large images | Downscale to about 448 pixels first |
| Long messages seem ignored in the middle (JEV-9B) | States over 1,024 tokens are trimmed (first 60%, last 40%) | Raise the limit, or use JEV-27B for long inputs |
nvidia-smi not found in WSL2 | Old or missing Windows NVIDIA driver | Update the driver on Windows, then restart WSL with wsl --shutdown |
| Results differ slightly between runs | bf16 math and batching | Normal in the third decimal place |
Easier alternatives for laptops and Macs
If none of the JEV models fit your hardware, you can still get Jev-style decisions locally, just from smaller models:
- Ollama decision models. Ollama 0.35 added Nimble and Tev1, served from a
/v1/systemoneendpoint that copies Jev's request shape. Tev1 0.8B is an 812 MB download that runs on almost anything. Our Ollama decision models guide walks through it. - llama.cpp with Kev-4B, Lev, Laya and Julia-1. Small, Apache 2.0 decision models with ready-made GGUF files that do include their decision parts, running on Windows, Linux and Macs.
- Cloudflare Clef and AWS Strands Decider 2B. Two more open decision models; Strands Decider runs on a CPU.
- Hosted Jev itself. At $0.042 per million input tokens, it is often the cheapest and simplest choice when privacy rules allow it.
These small models are less accurate than JEV-27B on hard questions, but for routing support email or tagging reviews, a 4B decision model on a laptop is often all a small business needs.
For IT admins: self-hosting JEV in a business
- License review is easy: Apache 2.0 across the family. Keep the license and notices with any redistribution. Steer staff away from OpenJev (CC BY-NC 4.0) for work use.
- Verify the artifact, not the name. Pin downloads to the official
autotrust/repositories and a specific revision, and check thathead.safetensors, the adapter folder andcalibration.jsonare present. A hollow GGUF on a shared drive will quietly turn a decision pipeline into a chatbot. - Keep the endpoint internal. vLLM's server has no login by default. Bind it to localhost or a private subnet, and front it with your existing gateway if other teams need access.
- Log the probabilities, not just the winner. Confidence is your audit trail; low-confidence decisions are exactly the ones a reviewer will ask about.
- Re-validate after every model update. These repositories changed several times in their first two weeks. Re-run your validation set before moving production to a new revision.
- Budget honestly. Below roughly a million decisions a day, hosted Jev usually costs less than a dedicated GPU; self-host for data control, not savings.
Open-source Jev models: frequently asked questions
Is the Jev model open source?
TypeSafe's Jev is closed and hosted only. AutoTrust's JEV-9B, JEV-27B, JEV-27B-VL and GEV-26B-Decide are open-weight copies of its decision behavior, released under Apache 2.0 on Hugging Face.
Is the Jev model free?
The JEV models are free to download and use commercially. Running them needs a capable NVIDIA GPU or a cloud instance. Hosted Jev costs $0.042 per million input tokens, with output free.
Where can I download the Jev model?
From Hugging Face, in the autotrust organization: autotrust/JEV-9B, autotrust/JEV-27B, autotrust/JEV-27B-VL, autotrust/GEV-26B-Decide and autotrust/GEV-26B-Decide-NVFP4. TypeSafe's own Jev has no download.
Is JEV-27B made by TypeSafe?
No. The JEV models come from AutoTrust AI, an independent group not affiliated with or endorsed by TypeSafe. They were trained on a public dataset of Jev's output distributions.
Why does my JEV GGUF chat instead of making decisions?
Most community GGUFs of JEV-9B, JEV-27B and GEV-26B-Decide contain only the base Qwen or Gemma weights. The decision adapter, head and calibration live in separate files that the conversion leaves out.
Can I run JEV models in Ollama?
Not usefully today. Ollama's library has no JEV models, and an imported hollow GGUF has no decision head for Ollama's /v1/systemone endpoint to read. Ollama's own Nimble and Tev1 are the working alternatives.
Can I run JEV models in LM Studio or llama.cpp?
Only Atlas3D's JEV-27B-VL GGUF makes decisions in llama.cpp, and it needs the patch shipped with it. In LM Studio, the hollow GGUFs load but behave like the plain base model.
How much VRAM does JEV-9B need?
Its full-precision weights are about 18 GB, so 24 GB of graphics memory is the practical floor for short decision prompts, and 32 to 48 GB is comfortable.
How much VRAM does JEV-27B-VL need?
AutoTrust's setup expects one GPU with 80 GB or more. Atlas3D's patched GGUF needs about 33.6 GB at Q8_0 or about 22.7 GB at Q4_K_M, with slightly different decisions at 4-bit.
Can I run the Jev model on a Mac?
There is no published Mac route for the JEV decision part, because its setup needs NVIDIA's CUDA. On a Mac, use small decision models through Ollama or llama.cpp instead.
Can I run JEV models on AWS?
Yes, on EC2. JEV-9B fits a g6.xlarge at about $0.80 an hour or a g6e.xlarge at about $1.86, and all four models fit a g7e.2xlarge with a 96 GB GPU at about $3.36 an hour in US East.
What is the difference between JEV-9B and JEV-27B?
JEV-9B is about 2.6 times faster and a third of the size. JEV-27B is closer to Jev, handles unfamiliar tasks and 16-option questions better, and scores higher on coding with its chat side.
What is GEV-26B-Decide?
A JEV-family model built on Gemma 4 26B-A4B that adds adaptive thinking: it answers instantly when confident and reasons first when not. It was first published as JEV-Gemma4-26B-A4B.
What are noul, choice and score?
The three Jev question types: noul is yes or no, choice picks one of your options, and score rates from 0 to 5. Each returns a probability for every possible answer.
What is the Jev model used for?
Routing support tickets, spam and moderation checks, tagging reviews, deciding whether a human should look at something, and judging screenshots or camera frames inside automated loops.
Is the Jev model better than an LLM?
It is a different tool. A decision model is faster, cheaper per answer and returns trustworthy probabilities, but it cannot write or explain. The JEV models include a normal chat side for that.
Is self-hosting JEV cheaper than the Jev API?
Usually not. A g6.xlarge running all month costs about the same as 14 billion tokens of hosted Jev, roughly 920,000 decisions a day at 500 tokens each. Self-host for privacy or control.
Does JEV-9B work with images?
Yes, since October 3, 2026. Its vision server runs the multimodal Qwen3.5-9B with the same decision adapter, and text decisions match the text-only model within 0.011.
Can I use JEV models commercially?
Yes. All four are Apache 2.0. OpenJev, a different 27B decision model, is CC BY-NC 4.0 and does not allow business use.
Jake deleted the 17 GB Gemma file without regret. It was never broken; it was just never a Jev. For now the shop uses hosted Jev, and when he is ready to keep customer messages in-house, he knows exactly which files to look for and which instance to rent. If a download ever left you wondering whether you did something wrong, you probably didn't. Sometimes the label and the box simply disagree, and now you know how to check the box.
📌 If you keep one line from this page
A Jev without its head is just the model it was built on.
Check for head.safetensors before you trust the name.
Revision note. Written October 6, 2026, as the JEV family passed a million downloads. If your download chatted when it should have decided, nothing was wrong with you or your PC; you were simply handed the wrong box.