OpenJev Locally: llama.cpp Now Runs Jev-Style Decision Models
Yes, you can now run a Jev-style decision model on your own machine with plain llama.cpp. On October 2, 2026, llama.cpp merged a new /v1/systemone endpoint, the same request shape as TypeSafe's hosted Jev API, along with ready-made GGUF files for five open decision models: OpenJev (27B, reads screenshots), Kev-4B, Lev, Laya and Julia-1. Jev itself stays closed and hosted-only. Here is the catch nobody puts in the headline: the model everyone is searching for, OpenJev, is licensed CC BY-NC 4.0, so you cannot use it for business at all, and its smallest file is 19 GB, which means a 24 GB graphics card. The two models that will actually run on a normal laptop, Kev-4B and Lev, are about 3 GB each, Apache 2.0, free for commercial use, and answer the exact same API.
Ethan's friend Priya sells handmade stationery online and drowns in support email every Monday: "where is my order", "can I change the color", "this arrived damaged", and a steady trickle of spam. Last month she wired up TypeSafe's hosted Jev to sort them and loved it, until she read that every email was leaving her laptop for someone else's server. On Friday Ethan forwarded her a headline, "llama.cpp now runs OpenJev locally", and she brought her laptop to Jake's shop that afternoon. "Can it do the Jev thing without the internet?" Jake opened the OpenJev model card, scrolled to the license line, and said, "Yes, but not with this one." That is where this page starts: what a decision model is, which of the five to run on which machine, the license trap, the exact install on Windows, Linux and a Mac, and the first request that sorts an email in under a second without a byte leaving the building.
If decision models are new to you, our guide to decision models in Ollama is the gentler first step: it explains the idea with Nimble and Tev1, which Ollama added at the end of September. This page is the next chapter. llama.cpp is the engine underneath Ollama, LM Studio and a dozen other tools, and now that it speaks the decision API natively, the whole open field of Jev-style models runs from one program on any operating system.
What is a decision model? Jev and "System One" in plain English
A normal chat model writes. You ask it something, and it produces words one at a time until it decides to stop. That is wonderful for explanations and terrible for software, because software wants an answer it can act on: this email is a refund request, this screenshot shows a login page, this review is negative. Getting that from a chat model means asking nicely, hoping it answers in the right format, and parsing whatever comes back.
A decision model skips the writing. You give it a situation, called the state, and one or more typed questions. It reads everything once and returns a probability for each possible answer. No sentences, no parsing, no "Sure! Here's the classification you asked for." Just numbers your code can trust. TypeSafe AI, which came out of stealth on September 15, 2026 with $40 million in seed funding, made the idea famous with Jev, and borrowed a name from psychology to describe it: Daniel Kahneman's "System 1" is fast, intuitive judgment, and "System 2" is slow, deliberate reasoning. Chat models are System 2. Jev is System One, and its API endpoint is literally named /v1/systemone.
The API asks three kinds of question:
- Choice: pick one option from a list you define, such as "refund", "shipping", "product question" or "spam". You get the winning choice plus a probability for every option and a confidence value.
- Noul: a yes-or-no question, answered with a single probability between 0 and 1. "Is this email angry?" 0.91 means almost certainly yes.
- Score: place something on a scale you describe in words, from 2 to 10 levels, such as "vague", "somewhat specific", "very specific". You get a score on that scale and the probability of each level.
Because the model reads the input once and answers in a single step, decisions come back in tens or hundreds of milliseconds, not the seconds a chat model needs to write a paragraph. That is why decision models suit jobs that happen thousands of times a day: sorting a support inbox, checking every form submission, letting a browser agent decide which button to click next. If the agent idea is new, this plain explanation of AI agents shows where a fast decision model fits inside one.
Can you run Jev locally? The honest answer
Not Jev itself. TypeSafe has not released Jev's weights, and its documentation describes it as a hosted service. "jev model download" and "jev model local" are common searches, and any file claiming to be Jev is something else. The hosted API costs $0.042 per million input tokens, with output not billed, and typically answers in about 100 milliseconds. It accepts text only: no images, audio or video. That pricing is cheap enough that many people will keep using it, and that is a perfectly sensible choice.
What changed is that you now have a real alternative when "hosted" is the problem rather than the price. Within days of Jev's launch, independent groups started training open models that answer the same API: Laya, Kev, Lev, Julia-1 and OpenJev among them. Until this week, running them meant each project's own server, Python scripts or a fork of llama.cpp. As of October 2, it means one mainstream program you may already have installed.
Here is how the three routes compare for someone deciding today.
| Hosted Jev | Ollama (Nimble, Tev1) | llama.cpp (OpenJev, Kev, Lev, Laya, Julia-1) | |
|---|---|---|---|
| Where it runs | TypeSafe's servers | Your machine | Your machine |
| Cost | $0.042 per million input tokens | Free, electricity only | Free, electricity only |
| Images | No, text only | No | OpenJev only, one image per request |
| Smallest download | None | 812 MB (Tev1 0.8B) | 168 MB (Julia-1) |
| Privacy | Your data leaves your machine | Stays local | Stays local |
| Best for | Top accuracy, no hardware | Beginners already using Ollama | Choice of five models, screenshots, any OS |
Priya's case lands squarely in the last column. Her emails contain customer names and addresses, she wants nothing leaving the laptop, and she does not need images. What she needed to hear next was which of the five to pick.
Which decision model should you run? OpenJev vs Kev-4B vs Lev vs Laya vs Julia-1
All five answer the same /v1/systemone API, so you can switch between them by changing one command. They differ in size, in what they were trained on, and, crucially, in license. These are the ggml-org GGUF builds that llama.cpp's new endpoint was tested against.
| Model | Size | Built on | File sizes | License | Runs well on |
|---|---|---|---|---|---|
| OpenJev | 27B | Qwen 3.5, vision | 19 GB (Q4_K_M), 28.6 GB (Q8_0), 53.8 GB (BF16) | CC BY-NC 4.0, non-commercial | 24 GB card; 32 GB Apple Silicon Mac |
| Kev-4B | 4B | Qwen 3.5 4B Base (by Jared Palmer) | 3.03 GB (Q4_K_M), 4.48 GB (Q8_0), 8.43 GB (BF16) | Apache 2.0 | Any 4 GB+ card, or a laptop CPU |
| Lev | 4B | Qwen 3.5 4B (by interfaze-ai) | 3.01 GB (Q4_K_M), 4.48 GB (Q8_0), 8.42 GB (BF16) | Apache 2.0 | Any 4 GB+ card, or a laptop CPU |
| Laya | 0.4B | ModernBERT (by Convai Innovations) | 449 MB (Q8_0), 844 MB (BF16) | Apache 2.0 | Any PC, CPU only is fine |
| Julia-1 | 0.1B | mmBERT-small, multilingual (by SupersonicLabs) | 168 MB (Q8_0), 303 MB (BF16) | Apache 2.0 | Anything, including a Raspberry Pi-class box |
How to read that table in practice:
OpenJev is the strongest and the only one that reads images, which makes it the one for screenshot work, such as a browser or desktop agent deciding what is on screen. Its makers report 84.0% on their 10,000-question text test against 85.4% for hosted Jev, and the same 39 of 100 browser tasks completed on the MiniWoB benchmark as Jev. Those are the makers' own numbers, but they are specific and the gap they admit to is small. The price is its size and its license.
Kev-4B and Lev are the practical middle. Both sit on Alibaba's Qwen 3.5 4B, both are about 3 GB at 4-bit, and both carry the Apache 2.0 license, which allows commercial use. Lev is described as a zero-shot classifier, meaning it is built to handle categories it has never seen in training, which is exactly what you want when your categories are your own ("refund", "custom order", "wholesale"). Kev-4B is the general text-decision model. If you can only try one, start with Kev-4B and test Lev against it on your own data.
Laya and Julia-1 are tiny encoder models, the same family as the BERT models that powered search engines for years, fine-tuned to answer the decision API. They run comfortably on a CPU, start instantly, and answer in milliseconds. They will be less accurate on subtle questions than a 4B model, but for clear-cut sorting, such as spam versus not spam or language detection, they are often enough. Julia-1 is built on a multilingual base, so it is the one to try if your inputs are not all in English.
Priya's laptop has an RTX 3050 with 4 GB of video memory. Jake put Kev-4B on it at 4-bit, with Laya as a fallback for when the graphics card was busy with something else.
The license trap: OpenJev is non-commercial
This is the section most people skip, and it is the one that can cost a business real money. The OpenJev weights are released under Creative Commons Attribution-NonCommercial 4.0. In plain terms: you may use, share and adapt the model, with credit, for non-commercial purposes only. Sorting your own personal email, research, learning, a hobby project: fine. Sorting customer email for a shop, powering a paid product, or running decisions inside a company's workflow: not allowed under that license, however small the business.
Two details make this easy to miss. First, the OpenJev repository mixes licenses: its helper and server code is Apache 2.0, which is commercial-friendly, while the model weights are CC BY-NC 4.0. People see "Apache 2.0" in the repository and assume it covers everything. It does not. Second, the base model it was trained from, Qwen, is Apache 2.0. A fine-tune can carry a stricter license than its base, and this one does.
The fix is simple, which is why it is worth knowing early. For any commercial use, run Kev-4B, Lev, Laya or Julia-1, all Apache 2.0 according to their model cards, or pay for hosted Jev. Use OpenJev for personal projects, research, or to measure how much accuracy you give up with a smaller model. Priya's shop is a business, so OpenJev was out from the start, which is what Jake meant by "not with this one". Kev-4B it was. If you are unsure whether your use counts as commercial, assume it does, or ask a lawyer before you build on it.
How to run OpenJev and Kev-4B with llama.cpp (Windows, Linux, Mac)
Everything here uses llama-server, the web-server program that ships with llama.cpp. You need a build from October 2, 2026 or later, because that is when the decision endpoint was merged. Any older copy will start the model and then answer /v1/systemone with "not found". If you have not used llama.cpp before, our comparison of local model tools explains where it fits beside Ollama and LM Studio.
Step 1: install or update llama.cpp
- Windows: open PowerShell and run
winget install llama.cpp, orwinget upgrade llama.cppif you already have it. Alternatively, download the latest release zip from the llama.cpp GitHub releases page. Pick the CUDA build for an NVIDIA card, Vulkan for AMD or Intel graphics, or the CPU build if you have no graphics card. - macOS and Linux:
brew install llama.cpp(orbrew upgrade llama.cpp). On Kali and other Linux systems without Homebrew, the release page has prebuilt Linux binaries, or you can build from source with CMake. - Check the version:
llama-server --version. The build should date from October 2, 2026 or later.
Step 2: start a decision model
The -hf flag downloads the model from Hugging Face on first run and caches it, so later starts are instant. Pick the line that matches your machine:
# laptop or any 4 GB+ graphics card (commercial use OK) llama-server -hf ggml-org/Kev-4B-GGUF:Q4_K_M llama-server -hf ggml-org/lev-GGUF:Q4_K_M # any PC, CPU only (commercial use OK) llama-server -hf ggml-org/Laya-GGUF:Q8_0 llama-server -hf ggml-org/Julia-1-GGUF:Q8_0 # 24 GB graphics card, non-commercial use only llama-server -hf ggml-org/OpenJev-GGUF:Q4_K_M
By default the server listens on http://127.0.0.1:8080. The server recognizes a decision model from metadata inside the GGUF file and switches its input and output handling automatically, so there is no special flag to remember. If you have a graphics card and the model is slow, add -ngl 99 to put every layer on the GPU. Newer builds do this automatically when the model fits. If port 8080 is taken, add --port 8081.
Step 3: ask your first question
With the server running, open a second terminal. This request asks Kev-4B to sort one customer email three ways at once: what kind of message it is, whether the customer is upset, and how urgent it is.
curl http://localhost:8080/v1/systemone -H "Content-Type: application/json" -d '{
"model": "kev-4b",
"state": "Hi, my order 4471 arrived today but the notebook cover is torn. I need it for a gift on Saturday. Can you send another one?",
"questions": {
"category": {
"type": "choice",
"instructions": "What is the customer asking for?",
"criteria": {
"damaged": "Item arrived broken or damaged",
"shipping": "Where is my order, delivery times",
"change": "Change or cancel an order",
"spam": "Not a real customer message"
}
},
"upset": { "type": "noul", "instructions": "Is the customer unhappy?" },
"urgency": {
"type": "score",
"instructions": "How time-sensitive is this?",
"criteria": ["No deadline", "Soon but flexible", "Hard deadline within a week"]
}
}
}'
The answer comes back as probabilities in Jev's answer format: a chosen category with a probability for every option and a confidence value, a single 0-to-1 number for the yes-or-no question, and a position on the urgency scale with the probability of each level. On Priya's laptop the whole round trip took a fraction of a second. On Windows PowerShell, the single quotes around the JSON can misbehave. Save the JSON to a file called q.json and use curl.exe http://localhost:8080/v1/systemone -H "Content-Type: application/json" --data-binary "@q.json" instead.
Apple Silicon: the MLX option for OpenJev
On a Mac, llama.cpp runs on the built-in GPU through Metal, so the commands above work as written. For OpenJev specifically, its makers also publish MLX builds for Apple Silicon at about 27 GB and 15 GB. The smaller one is the realistic choice for a 32 GB Mac, which needs room left over for macOS and your apps. A 16 GB Mac should stay with Kev-4B or Lev.
The /v1/systemone request, explained field by field
Every Jev-style model on this page reads the same request shape, so once you understand one request you can use any of them, or the hosted API, without rewriting your code.
state is the situation being judged. It can be a plain string, like the email above, or a JSON object with named fields, such as {"name": "...", "description": "..."}. Your instructions can then point at a specific field. For OpenJev, the state can include one image, such as a screenshot.
questions is a set of named questions, each answered independently in the same pass. Names are yours to choose. They come back as keys in the answer, so pick names your code will read easily.
type is choice, noul or score, as described earlier.
instructions is the question in plain English. Short and specific beats long and clever. "Is the customer unhappy?" works better than a paragraph about tone.
criteria describes the options. For a choice question it is a map from an option name to a description. For a score question it is a list of level descriptions, lowest first. Noul questions have no criteria. The descriptions matter more than the names: the model reads "Item arrived broken or damaged", not just "damaged".
A few limits are worth knowing before you design around them. Hosted Jev allows up to 255 options in one choice question and 2 to 10 score levels, with the state and longest question fitting in about 32,000 tokens. OpenJev's own server handles up to 52 options in a single pass and splits bigger sets into several passes, with a 16,384-token prompt limit and one image per request. The small encoder models, Laya and Julia-1, have much shorter input limits than the 4B and 27B models, so long documents belong with Kev, Lev or OpenJev.
Using a local decision model from Python
You do not need a special SDK. Any HTTP library works, and Python's built-in one is enough. This script sorts every email in a folder of text files and prints the category with its probability.
import json, pathlib, urllib.request
URL = "http://localhost:8080/v1/systemone"
QUESTIONS = {
"category": {
"type": "choice",
"instructions": "What is the customer asking for?",
"criteria": {
"damaged": "Item arrived broken or damaged",
"shipping": "Where is my order, delivery times",
"change": "Change or cancel an order",
"spam": "Not a real customer message",
},
},
"upset": {"type": "noul", "instructions": "Is the customer unhappy?"},
}
def decide(text):
body = json.dumps({"model": "local", "state": text, "questions": QUESTIONS}).encode()
req = urllib.request.Request(URL, data=body, headers={"Content-Type": "application/json"})
return json.loads(urllib.request.urlopen(req, timeout=30).read())
for f in sorted(pathlib.Path("inbox").glob("*.txt")):
a = decide(f.read_text(encoding="utf-8"))["answers"]
print(f.name, a["category"]["choice"], round(a["upset"]["noul"], 2))
Two habits make this production-worthy. First, use the probabilities, not just the winning choice. If the top category is below, say, 0.6, send the email to a human instead of acting on it automatically. That single rule catches most of the mistakes a small model makes. Second, keep a sample of a few hundred real messages with the correct answers written down, and rerun them whenever you switch models. That labeled set is how Priya chose between Kev-4B and Lev, and it is the only benchmark that matters for your own data.
How good are open decision models compared with Jev?
Honest answer: close at the top, a real step down at the bottom, and the only test that counts is your own data.
OpenJev's makers publish the most detailed comparison. On their 10,000 text questions from 34 public sources, OpenJev scores 84.0% to hosted Jev's 85.4%, a gap of 1.4 points. On 100 MiniWoB browser tasks both complete 39. They also report a stability number that matters more than it sounds: when the same options are shuffled into a different order, OpenJev's answer changes 2.3% of the time, against 18.5% for the untuned base model. A model whose answer depends on the order of your list is a model you cannot trust, and that is exactly what decision-model training fixes.
For the smaller models, the ggml-org model cards publish no benchmark numbers, and independent leaderboards are still catching up with how fast this field moves. One community benchmark site places a smaller project that also calls itself OpenJev well below Jev, which shows how quickly names and numbers get mixed up in a field this young. The practical rule: expect a 4B model to handle clear categories well and to struggle more with subtle ones, such as sarcasm or mixed requests, and expect the sub-1B encoders to be best at simple, high-volume sorting.
Calibration deserves a warning of its own. A decision model's probabilities are only useful if 0.9 really means "right about nine times in ten". OpenJev ships with fixed calibration settings for exactly this reason. Some third-party checkpoints have shipped with badly tuned settings, so before you build a threshold like "act automatically above 0.8", check on your labeled sample that answers above 0.8 really are right that often. If they are not, raise the threshold until they are.
"openjev github": telling the OpenJev projects apart
"openjev github" is the top search for this model, and the results are confusing for a simple reason: several unrelated projects use nearly the same name.
openjev/openjevon Hugging Face: the 27B Qwen-based decision model with vision, CC BY-NC 4.0 weights. This is the one llama.cpp supports, converted by ggml-org asggml-org/OpenJev-GGUF. It states plainly that it is an independent project, not affiliated with TypeSafe.- A llama.cpp fork named openjev.cpp: built for a different, much smaller openjev cross-encoder. Useful history, but you no longer need a fork, because mainline llama.cpp now has the decision endpoint.
- Docker server projects named openjev: community packaging that runs a Jev-compatible API on an NVIDIA card. Fine if you prefer containers, but check which weights they download and what license those carry.
- Smaller "open Jev" replicas: research projects of a few hundred million parameters, some MIT-licensed, built to show how a Jev-style model is trained.
The rule that keeps you out of trouble: start from the model name in the llama-server -hf command, ggml-org/OpenJev-GGUF, and follow its "base model" link to the original. Do not assume a project's license or benchmark from its name.
Can Ollama or LM Studio run OpenJev?
Partly, and this is where many first attempts go wrong. The ggml-org model page offers a one-click "use this model" button for Ollama, LM Studio, Jan and others, and Ollama will happily download the file with ollama run hf.co/ggml-org/OpenJev-GGUF:Q4_K_M. Downloading is not the same as answering decisions, though. Ollama added its own /v1/systemone endpoint in version 0.35.0 at the end of September for its own decision models, Nimble and Tev1. Whether that endpoint accepts these llama.cpp-converted files is not documented yet. If you load one as an ordinary chat model, it will not behave like a decision model at all.
Until that is settled, the reliable split is:
- Decision models from the Ollama library (Nimble, Tev1): use Ollama, as in our Ollama decision-models guide.
- The five ggml-org decision GGUFs (OpenJev, Kev-4B, Lev, Laya, Julia-1): use llama-server, as on this page.
- Both at once: fine. Ollama uses port 11434 and llama-server uses 8080, so they do not collide, and you can compare models from the two worlds on the same labeled sample.
LM Studio bundles its own copy of llama.cpp, so decision support will reach it when its bundled engine updates past October 2. Until then, the standalone llama-server is the way.
What hardware do you need for OpenJev?
For the 27B OpenJev, the hard number is the file size: 19 GB at 4-bit. Add working memory for the input, especially with an image, and a 24 GB graphics card such as an RTX 3090 or 4090 is the realistic minimum for running it fully on the GPU. The 8-bit file at 28.6 GB needs a 32 GB card or two cards. The full-precision 53.8 GB file is for servers. Its makers' reference setup is a single 80 GB H100 at FP8, where they measure about 80 milliseconds for a short text decision and 176 milliseconds for a desktop screenshot.
On a machine with less video memory, llama.cpp can split the model between the graphics card and system memory. It works, but every decision then takes seconds rather than milliseconds, which throws away the main point of a decision model. If OpenJev does not fit your card, a 4B model fully on the GPU will almost always serve you better than a 27B model half on the CPU. The honest laptop list covers which machines carry 24 GB if you are shopping.
The small models are the opposite story. Kev-4B and Lev at 4-bit fit comfortably in 4 GB of video memory, and run acceptably on a modern laptop CPU. Laya and Julia-1 need almost nothing: under 1 GB of memory, no graphics card, and answers fast enough for thousands of decisions a minute on an ordinary desktop.
Where Cloudflare Clef fits
The day before llama.cpp's merge, Cloudflare released Clef and Clef-flash, two decision models that speak the same request shape and can read up to four images, hosted on Workers AI with a free daily allowance. Our Clef guide covers the pricing and setup. On the local question, Clef is not among the five models in llama.cpp's decision collection, and the community GGUF repacks posted on launch day carry the language backbone but not the part that makes the decision. So for now, Clef is a hosted option and OpenJev is the local route for image decisions. Both pages will get a dated line when that changes.
That leaves a clean map for anyone choosing today. Hosted Jev for the best text accuracy with no hardware. Clef-flash for hosted decisions about photos with a free tier. Ollama's Nimble and Tev1 if you already live in Ollama. llama.cpp with Kev-4B or Lev for private, commercial-friendly local decisions on a laptop. OpenJev on a 24 GB card for local screenshot decisions in non-commercial work.
What can you do with a local decision model? Six real jobs
"How to use an LLM for text classification" is one of the most searched questions in this whole area, and a decision model is the most direct answer to it. Here are six jobs where a small local model earns its place, with the question type and model size that suit each one.
| Job | Question to ask | Type | Start with |
|---|---|---|---|
| Support inbox sorting | Refund, shipping, change, or spam? | choice + noul for "upset" | Kev-4B or Lev |
| Contact-form spam filter | Is this a real person with a real question? | noul | Laya |
| Review and comment moderation | Positive, negative, abusive, or off-topic? | choice | Kev-4B; Julia-1 for many languages |
| Security log triage | How suspicious is this log line? | score | Lev, run on the Kali box itself |
| Document routing | Invoice, contract, receipt, or letter? | choice | Kev-4B (long inputs) |
| Browser or desktop agent steps | Which button on this screenshot comes next? | choice with an image | OpenJev (non-commercial) |
The security row deserves a word for anyone running Kali. Log triage is a classic decision job: thousands of lines, most of them boring, a few that matter. A score question such as "How suspicious is this authentication event?" with three or four described levels lets a 4B model running on the same machine sort a night of logs before you open them, without shipping a single line to an outside service. Treat its output as a ranking for your attention, not a verdict. The model has never seen your network, and attackers do not label their traffic.
Across all six, the pattern that works is the same. Define a small number of options with clear descriptions, ask one question per decision, act automatically only above a confidence threshold you have checked, and route everything else to a person. Decision models are at their best as a very fast first pass, which is exactly the role that used to need either a trained classifier or a slow and expensive chat model.
Decision model vs BERT classifier vs prompting a chat model
Before decision models, there were two ways to classify text with AI, and both are still around. Knowing why decision models beat them for most small projects helps you choose, and also spot the cases where an older method is still right.
A trained BERT-style classifier is the traditional approach: take a small encoder model and train it on a few thousand labeled examples of your exact categories. It is fast, tiny and very accurate on those categories. The catch is the training data. You need the labeled examples before you start, and every new category means more labeling and retraining. Laya and Julia-1 are, interestingly, built from exactly this family of models. The difference is that they were trained to follow the instructions and criteria you write, so you get a BERT-sized model without training your own.
Prompting a chat model is the newer habit: paste the text into a chat model with "classify this as A, B or C" and read the reply. It needs no training and understands nuance well. It is also slow, because the model writes its answer word by word. It is inconsistent, because the same email can get "A" one time and "Category: A." the next. And it gives no honest probability, because a chat model's stated confidence is just more generated text.
A decision model takes the best of both. Like a chat model, it needs no training: you describe the categories in words. Like a classifier, it answers in one pass with real probabilities. The trade-off is the one-off decisions where the reasoning matters more than the label, such as "explain why this contract clause is risky". Those still belong to a chat model, so use the decision model to decide what deserves a closer look, and a chat model to give that closer look.
Common errors and fixes
- "404 Not Found" on /v1/systemone. Your llama.cpp is older than October 2, 2026. Update it and check
llama-server --version. - The server starts, then answers like a chat model. You loaded a normal GGUF, not one of the ggml-org decision builds. Use the exact
-hf ggml-org/...-GGUFnames above. The decision behavior comes from metadata in those files. - Out of memory loading OpenJev. The 19 GB file does not fit your card. Use Kev-4B, or lower the GPU layers with
-ngland accept slow answers. - Every answer is slow. The model is running on the CPU. Add
-ngl 99, check that you installed the CUDA or Vulkan build rather than the CPU build, and make sure no other program is filling video memory. - PowerShell says the JSON is invalid. PowerShell rewrites quotes in arguments. Save the request to a file and send it with
curl.exe --data-binary "@q.json". - "Address already in use". Something else is on port 8080. Start with
--port 8081and change the URL to match. - Probabilities look overconfident. Check calibration on your labeled sample before trusting any threshold. Some community checkpoints ship with poorly tuned settings.
Keeping a local decision server private
llama-server binds to 127.0.0.1 by default, so only your own machine can reach it. Keep it that way unless you need it on your network. If you do open it with --host 0.0.0.0, add --api-key with a long random value and firewall the port, because without a key anyone who can reach it can use your graphics card and read whatever your programs send it. Never forward the port from your router to the internet. For access from elsewhere, put it behind a VPN or a private network layer.
The privacy upside is the reason most people are here: nothing in a state leaves your machine. For customer email, medical notes, HR documents or anything else you would hesitate to paste into a website, that is the whole argument. It is also worth remembering that a decision model returns a judgment without an explanation. Use it to sort, flag and route, and keep a human on anything with real consequences for a person.
OpenJev, Jev and local decision models: frequently asked questions
What is OpenJev?
OpenJev is an open 27-billion-parameter decision model built on Qwen 3.5 that answers TypeSafe's Jev API format, including one image per request. It is an independent project, not affiliated with TypeSafe, and its weights are licensed CC BY-NC 4.0 for non-commercial use.
Can you run Jev locally?
No. TypeSafe has not released Jev's weights; it is a hosted API. Open Jev-style models such as OpenJev, Kev-4B, Lev, Laya and Julia-1 run locally in llama.cpp through the same /v1/systemone request shape.
How do I run OpenJev with llama.cpp?
Update llama.cpp to a build from October 2, 2026 or later, then run llama-server -hf ggml-org/OpenJev-GGUF:Q4_K_M and send requests to http://localhost:8080/v1/systemone. You need about 24 GB of video memory for the 19 GB file.
Is OpenJev free for commercial use?
No. OpenJev's weights are licensed CC BY-NC 4.0, which allows non-commercial use only; its helper and server code is Apache 2.0. For commercial work, use Kev-4B, Lev, Laya or Julia-1, which are Apache 2.0, or the hosted Jev API.
Where is OpenJev on GitHub and Hugging Face?
The model is on Hugging Face as openjev/openjev, and the llama.cpp-ready files are at ggml-org/OpenJev-GGUF. Several unrelated GitHub projects use similar names, so start from the ggml-org model and follow its base-model link.
How does OpenJev compare with Jev in benchmarks?
Its makers report 84.0% on their 10,000-question text test against 85.4% for hosted Jev, and the same 39 of 100 MiniWoB browser tasks. Answer changes from reordering the options drop to 2.3%, from 18.5% for the untuned base model.
What is the /v1/systemone API?
It is the decision-model endpoint introduced by TypeSafe's Jev. You send a state and typed questions (choice, noul or score), and get probabilities back instead of generated text. llama.cpp added the same endpoint on October 2, 2026, and Ollama has one for its own decision models.
What is a noul question?
A noul question is a yes-or-no question in the Jev API. It returns a single probability between 0 and 1, where values near 1 mean yes and values near 0 mean no, with no separate confidence field.
Which decision model is best for a laptop?
Kev-4B or Lev at 4-bit, about 3 GB each, which fit a 4 GB graphics card or run on a modern laptop CPU. Both are Apache 2.0. For an older PC with no graphics card, Laya (449 MB) or Julia-1 (168 MB) run comfortably on the CPU.
What GPU do I need for OpenJev?
A 24 GB card such as an RTX 3090 or 4090 for the 19 GB 4-bit file, or a 32 GB Apple Silicon Mac with the smaller MLX build. The 28.6 GB 8-bit file needs more memory, and the full 53.8 GB model is meant for data-center GPUs.
Can Ollama run OpenJev as a decision model?
Ollama can download the GGUF, but whether its /v1/systemone endpoint accepts the ggml-org decision files is not documented yet. Use llama-server for OpenJev, Kev-4B, Lev, Laya and Julia-1, and Ollama for its own Nimble and Tev1.
Does OpenJev support images?
Yes, one image per request, such as a desktop or web screenshot, which makes it suited to browser and desktop agents. Kev-4B, Lev, Laya and Julia-1 are text only, and hosted Jev currently accepts text only.
How much does the Jev API cost?
Hosted Jev costs $0.042 per million input tokens, with output not billed, and typically answers in about 100 milliseconds. Running open decision models locally costs nothing beyond electricity.
What is the difference between a decision model and an LLM?
A chat LLM generates text one token at a time. A decision model reads the input once and returns a probability for each option you define, so answers are typed, fast and need no parsing. Many decision models are built from LLMs, with the text generation replaced by a decision head.
Is Jev open source?
No. Jev is a closed, hosted model from TypeSafe AI. The open alternatives are separate projects trained to answer the same API, such as OpenJev, Kev-4B, Lev, Laya and Julia-1.
Why does llama.cpp say 404 on /v1/systemone?
Your llama.cpp build is older than October 2, 2026, when the decision endpoint was merged. Update with winget, Homebrew or the latest release zip, and confirm with llama-server --version.
Priya's Monday inbox looked different the following week. Kev-4B sorted 212 emails before her coffee was finished. Damaged items went to one folder with a red flag, where-is-my-order questions to a folder she answers with a template, and the 31 spam messages straight to the bin. Anything the model was less than 60% sure about, 14 emails, waited for her in a folder called "you decide". Nothing left the laptop. She sent Jake a photo of the folder counts with a single line: "It's like having a very fast intern who never reads my customers' addresses out loud." That is what a decision model is for: not wisdom, but fast, honest sorting, with the hard calls left to a person. And as of this week, it runs anywhere llama.cpp does.
📌 If you keep one line from this page
OpenJev is the headline, but Kev-4B is the one most of us should run, and it is free for business.
Update llama.cpp, run llama-server -hf ggml-org/Kev-4B-GGUF:Q4_K_M, and send anything under 0.6 confidence to a human.
Revision note. Written October 2, 2026, the day llama.cpp merged its decision endpoint. If your inbox has been eating your Mondays, I hope this gives one of them back.