Cloudflare Clef: Pricing, Clef vs Jev, and Can You Run It Locally?

Logeshwaran
—
Cloudflare Clef: Pricing, Clef vs Jev, and Can You Run It Locally?

Cloudflare Clef is a pair of decision models, released October 1, 2026, that answer typed questions about a piece of text or an image with a probability for every option instead of writing a reply. Clef is 27 billion parameters; Clef-flash is 9 billion and answers in a median of 38.8 milliseconds on Cloudflare's own tests. Both sit in the Cloudflare Workers AI models catalog as @cf/cloudflare/clef and @cf/cloudflare/clef-flash, both are Apache 2.0 with the weights on Hugging Face, and both speak the same request shape as TypeSafe's Jev and Ollama's /v1/systemone. Clef costs $0.24 per million input tokens and Clef-flash $0.09, which is 5.7 and 2.1 times Jev's price per token, and yet for a small business it is usually free, because Workers AI gives every account 10,000 Neurons a day and that buys about 1.2 million Clef-flash tokens. Here is the twist nobody is leading with: the "open source" model you can download in 19 GB, which the community re-quantized into GGUF and MLX files within four hours of launch, cannot be asked a single question on a laptop today. The part that makes the decision is a separate 244 MB file that no desktop runtime loads yet. The second twist is smaller and stranger: on three of Cloudflare's own headline benchmarks, the 9B model beats the 27B one.

Jake runs a phone-repair shop. The shop PC has been filing his 300 weekly messages with a local decision model since the Sunday the Ollama decision models arrived, and it works. What it cannot do is look at the photos. Half his customers send a picture of the damage before they send a sentence, and a cracked screen, a swollen battery and a phone that went through the wash all need different replies. Ethan, his mentor, read the Clef announcement over his shoulder and said: "You already have a clerk in the shop. This one works at the post office down the road, reads photographs, and the first twelve hundred thousand words a day are free. You do not have to choose." The rest of this page is Jake's questions, which are probably yours.

⚡ Quick Answer

• What it is → two Cloudflare-trained decision models: Clef (27B, built on Qwen3.8-27B) and Clef-flash (9B, on Qwen3.5-9B). Text plus up to four images in, a choice, a yes/no probability or a score out. Details.

• Price → $0.24 per million input tokens for Clef, $0.09 for Clef-flash, output never billed. 10,000 free Neurons a day on every Workers AI account, about 1.2M Clef-flash tokens or 458K Clef tokens. The math.

• First call → env.AI.run("@cf/cloudflare/clef-flash", {...}) in a Worker, or one curl to api.cloudflare.com/client/v4/accounts/ID/ai/run/@cf/cloudflare/clef-flash. The exact JSON.

• Locally? → weights yes, laptop no. Clef-flash wants about 41 GB of VRAM at full precision; the GGUF files are missing the decision head; Ollama has no tag. The honest answer and what to run instead.

Clef vs Jev: 2.5 to 13 times faster, 2 to 6 times the token price, and the "beats Jev" claim is Cloudflare's own measurement, not yet on the public leaderboard. The comparison.

Cloudflare Clef and Clef-flash, released October 1, 2026: open Apache 2.0 decision models on Workers AI, $0.09 and $0.24 per million input tokens with 10,000 free Neurons a day (about 1.2 million Clef-flash tokens), a 38.8 ms median decision for Clef-flash, up to four images per request, and GGUF and MLX repacks that lack the 244 MB decision head so nothing on a laptop can run it yet

If "decision model" is new to you, the plain-English guide to Ollama's decision models is the chapter before this one and explains the request shape slowly; the guide to what an LLM is is the chapter before that. You do not need either to use Clef, but the trap section near the end assumes you have met the idea that a model which cannot answer outside your list can still answer wrong.

What Cloudflare Clef is, in plain English

A chat model writes. You give it a prompt and it produces an answer one token at a time, which is why it can explain itself, and why it can also hand you a category you never offered, a trailing sentence after the JSON, or a polite refusal. A decision model does not write. You give it a state, the text or image you are asking about, and a set of typed questions, each with the options you will accept. It scores every option in one forward pass and returns the whole probability distribution. It cannot add an option. It cannot say "possibly both." It cannot explain, either, which is the price you pay for the other two.

Clef is Cloudflare's version of that idea, and the first model family its Workers AI team has trained rather than hosted. Underneath, Clef is Qwen3.8-27B with its vision encoder, frozen, with a joint schema head trained on top; Clef-flash is the same recipe on Qwen3.5-9B. The head is what turns a language model into a decision model: instead of predicting the next token, it emits one logit per allowed option for every question, and a softmax per question turns those into probabilities. Cloudflare describes the architecture as non-autoregressive and prefill-only, with "two-stage attention routing with cross-field attention" so that several questions about the same state are scored in parallel rather than one after another. The training used label-smoothed cross-entropy, a Brier loss to keep the probabilities honest, and a reinforcement-learning stage they call RLCD, Reinforcement Learning for Calibrated Decisions.

The request shape is not Cloudflare's. It is the one TypeSafe published for Jev in mid-September 2026 and that Ollama copied into /v1/systemone two weeks later: model, state, questions. Clef accepts the same three question types, noul for a yes/no answered as a probability, choice for one option from a named list, and score for a position on an ordered scale. If you have written a request for Nimble or Tev1, you have written one for Clef. What Clef adds is a vision encoder, so the state can include up to four images, and a 64K-token context on the hosted service.

  Clef Clef-flash
Size and base27B, Qwen3.8-27B with vision encoder9B, Qwen3.5-9B with vision encoder
Workers AI model ID@cf/cloudflare/clef@cf/cloudflare/clef-flash
Price$0.24 per M input tokens (21,818 Neurons per M); output free$0.09 per M input tokens (8,182 Neurons per M); output free
Median latency, Cloudflare's 43-benchmark run209.3 ms (p95 238.6 ms)38.8 ms (p95 122.4 ms)
Context65,536 tokens hosted; trained for 256K65,536 tokens hosted; trained for 256K
ImagesUp to 4 per requestUp to 4 per request
WeightsCloudflare/clef, 55 GB, 12 shards plus a 256 MB joint headCloudflare/clef-flash, 19.1 GB, 4 shards plus a 244 MB joint head
LicenseApache 2.0Apache 2.0
Training dataNot published. Cloudflare's product manager for the AI platform confirmed the datasets are private, which is worth knowing before you repeat "open source" in a procurement form.

Jake: "So it is Nimble with eyes, hosted by the company that already serves my website."

Ethan: "Close. Nimble is a LoRA on a 9B Qwen; Clef-flash is a full head on the same size of Qwen with the camera left on, trained by a team with a lot of GPUs and a lot of labeled web traffic. Same request, same three question types, same inability to make things up. The differences are the pictures, the speed, and where it runs. Keep the shop PC doing what it does. Add this for the photographs."

🧭 NEW HERE? READ THESE FIRST

New to decision models, or to the words on this page? These explain the pieces without the jargon:

 Bookmark this page; it is the hosted-decision-model chapter of the run-AI-locally series, and the one where "local" needs an honest asterisk.

Clef vs Clef-flash: the small one wins three of the headline tests

Every model launch has a big one and a fast one, and the big one is supposed to be better. Read Cloudflare's own numbers slowly and that story breaks. On BFCL, the function-calling benchmark, Clef-flash scores 98.76% exact-case against Clef's 98.47%. On API-Bank, Clef-flash reaches 93.11% accuracy. On the TypeSafe customer-service workflow evaluation, Clef-flash posts 77% where Clef is the model Cloudflare quotes for the harder invoice-processing (64.7%) and security-incident (62.9%) sets. The large model wins where the question needs more reading: ToolRet retrieval at 69.19% nDCG@10, BANKING77 intent classification at 94.20% macro-F1, CLINC150 with out-of-scope detection at 97.43%.

That pattern is not a mistake and it is not marketing. Decision models are scored on bounded questions, and a 9B model that has been trained to be calibrated on short, well-defined choices can beat a 27B one on exactly those choices, while losing when the state is long, the options are many, or the right answer depends on world knowledge. The practical rule falls out of it: start with Clef-flash. It is a third of the price, five times faster, and on the kind of question most businesses actually ask, "which team," "is this urgent," "is this spam," it is as good or better. Move a question to Clef only when your own test set says the small one is getting it wrong.

Benchmark (Cloudflare's published run) Clef 27B Clef-flash 9B What it measures
BFCL, case exact98.47%98.76%Picking the right tool and arguments for an agent
API-Bank, accuracynot quoted93.11%Deciding which API call a request needs
Typesafe Workflowevals, customer servicenot quoted77%Multi-step support triage
Typesafe Workflowevals, invoice processing64.7%not quotedReading fields and routing documents
Typesafe Workflowevals, security incidents62.9%not quotedSeverity and routing of alerts
ToolRet, nDCG@1069.19%not quotedRanking the right tool from a long list
BANKING77, macro-F194.20%not quoted77-way intent classification
CLINC150 + out-of-scope, macro-F197.43%not quoted150 intents plus "none of these"
Median latency across 43 benchmarks209.3 ms38.8 msOn Cloudflare's GPUs, not including your network

One honest note on the "not quoted" cells: Cloudflare's launch post picks the best model for each row rather than printing both, so the table above is what they chose to show, not a full grid. When the complete Decision Index run for both models is published you will be able to fill the gaps; until then, treat the rows as "the better of the two" and test your own question on both. It costs a fraction of a cent.

Clef vs Jev: faster, pricier per token, and a claim that is still Cloudflare's own

Jev is TypeSafe's hosted decision model, the one that created the category in mid-September 2026 and whose API everybody else copied. It is a paid, early-access service at $0.042 per million input tokens, with output free and no downloadable weights. Clef is the first serious challenge to it from a company with its own global GPU fleet, and Cloudflare's launch material is built around the comparison: Clef-flash answers 13 times faster than Jev, Clef 2.5 times faster, and Clef "leads the evals" on the Jev Decision Index, winning 7 of 10 decision benchmarks.

Three things sit behind those sentences and you should hold all three at once. First, the speed numbers are real but not symmetrical: Clef's 38.8 and 209.3 milliseconds were measured on Cloudflare's own cards, while Jev's 524.1 milliseconds includes a network round trip to TypeSafe's API from a lab, and the public index page says so in its notes. From a developer's machine in Italy on launch day, the same REST call to Clef-flash took 191 to 205 milliseconds and Clef 524 to 726, which is still faster than Jev from the same place, just not thirteen times. Second, the price runs the other way: per token, Clef costs about 5.7 times Jev and Clef-flash about 2.1 times. For a workload of thousands of short decisions a day that gap is pennies, and the free tier erases it entirely below a threshold we will calculate; for a workload of millions of long documents it is real money, and Jev is cheaper. Third, the accuracy claim. Cloudflare says Clef beat Jev in three of four areas on TypeSafe's own benchmark set, losing only on agent-trace observability, and posts an index score of 61.2 for Clef against 57.9 for Jev and 57.1 for Clef-flash. Those are Cloudflare's measurements on the public suite, not an entry on the independent leaderboard; the public results file, generated September 28, has Jev at 57.91 and no Clef row yet. Nobody is calling the number wrong. It is simply unverified, and a careful reader keeps "self-reported" pinned to it until the leaderboard refreshes.

  Jev (TypeSafe) Clef Clef-flash
Price per M input tokens$0.042$0.24 (5.7x)$0.09 (2.1x)
Free tierEarly access; none published10,000 Neurons a day on every Workers AI account
Median latency524.1 ms (hosted, includes network)209.3 ms on-card; 524 to 726 ms from Italy38.8 ms on-card; 191 to 205 ms from Italy
Decision Index score57.91 (public results, Sept 28)61.2 (Cloudflare's own run)57.1 (Cloudflare's own run)
Reasoning-heavy questionsStronger: GPQA Diamond 78.3Weaker: GPQA Diamond 48.0Weaker still
ImagesNoUp to 4Up to 4
Context64K per request; state plus longest question capped at 32K64K64K
WeightsClosed; architecture undisclosedApache 2.0, 55 GBApache 2.0, 19.1 GB
Run it yourselfNeverOn an 85 GB GPU, with Cloudflare's PythonOn a 41 GB GPU, with Cloudflare's Python

The GPQA Diamond row is the one that tells you what kind of model Clef is. Jev scores 78.3 on graduate-level science questions and Clef 48.0, and that is fine, because nobody should be using a decision model to answer graduate-level science questions. Clef wins on classification, routing and tool choice, loses on reasoning, and costs more per token while being faster. If your question is "which of these twelve labels," Clef-flash is the better bet today. If your question needs the model to think before it chooses, neither of these is the right tool, and the big hosted models remain the place to send that case.

Clef vs Nimble, Tev1, Laya and the rest: where it sits on the Decision Index

The Jev Decision Index is the public scoreboard for this whole category: a frozen suite of 120,340 requests across 43 benchmarks, 38 of them in the scored panel, five equal-weight areas (Tools and Automation, Retrieval and Classification, Language Understanding, Knowledge and Reasoning, Arts and Human Judgment), run on one RTX PRO 6000 with house rules that forbid truncating requests, dropping options or tuning prompts per model. Its headline figure is a chance-normalized "skill" score, so a model that guesses randomly scores near zero rather than near 25%. That is why the numbers below look lower than the accuracy figures on model cards; they are measuring how much better than chance a model is across everything, not how often it is right on one friendly dataset.

Model Size, base Index skill score Median latency Where it runs
Clef27B, Qwen3.8-27B61.2 (Cloudflare's run)209.3 msWorkers AI; 85 GB GPU
Jev 1.13Undisclosed57.91524.1 ms (network included)TypeSafe API only
Surogate Rune 26B-A4B v326B MoE, Gemma 457.44120.5 msOpen weights, big GPU
Clef-flash9B, Qwen3.5-9B57.1 (Cloudflare's run)38.8 msWorkers AI; 41 GB GPU
AutoJev-27B27B, Qwen3.8-27B56.4101.4 msOpen weights, big GPU
Decider 4B4B, Qwen3.5-4B40.712.6 msOpen weights
Winnow-E4B8B effective, Gemma 4 E4B39.8945.0 msOpen weights; ollaya
Nimble 9B v29B, Qwen3.5-9B39.5777.3 msollama pull nimble
lev4B, Qwen3.5-4B38.5470.8 msGGUF on ggml-org since Oct 1
Kev 9B9B, Qwen3.5-9B38.4851.4 msGGUF on ggml-org (4B) since Oct 1
Tev1 4B4B, Qwen3.5-4B29.2435.8 msollama pull tev1
Decider 2B2B, Qwen3.5-2B28.978.1 msOpen weights
Tev1 0.8B0.8B, Qwen3.5-0.8B12.8528.5 msollama pull tev1:0.8b
Laya421M, ModernBERT-large encoder6.045.8 mspip install laya; GGUF on ggml-org
Julia 1141M, mmBERT-small encoder5.545.8 msGGUF on ggml-org

Read down that table and two lessons jump out. The gap between Clef-flash at 57.1 and Nimble at 39.57 is not a rounding error; it is the difference between a full schema head trained with reinforcement learning on a large private set and a rank-16 LoRA trained on 2,676 curated examples. The second lesson is about the tiny encoder models at the bottom. Laya's own model card reports 76.6% on a typed-decisions set and a 32.8 ms median, both better than Jev on the card's own comparison, and those numbers are honest for the narrow set they were measured on. On the full 38-benchmark panel with no prompt tuning, Laya scores 6.04, near chance. That is the single most useful thing the index teaches: a decision model can be excellent at the questions it was built for and useless one benchmark over, so the test that matters is yours.

‍♂️ Jake's Reality Check

"Nimble scored 75 percent on the Ollama page last week and 39 here. Which one is lying?"

The straight answer. Neither. The Ollama figure is plain accuracy on 13 datasets; the index figure is how far above random guessing a model lands across 38 benchmarks, where a coin flip scores zero. Same model, two rulers. On your 300 messages a week, Tev1 4B has been right often enough that you stopped checking, and that is the only ruler that pays your rent. Use the index to choose which models to try, and your own 50 messages to choose which one to keep.

Cloudflare Workers AI pricing for Clef: the free tier limits and the math

Workers AI bills in Neurons, a unit that lets Cloudflare price very different models on one meter. Every account gets 10,000 Neurons a day at no charge, on the free plan and the paid one alike; past that, the Workers Paid plan charges $0.011 per 1,000 Neurons. Clef is listed at 21,818 Neurons per million input tokens, which is where the $0.24 comes from, and Clef-flash at 8,182 Neurons per million, which is the $0.09. Output tokens are never counted, because a decision model emits none; the usage block in every response reports output as zero.

Turn the free allocation into words. 10,000 Neurons divided by 8,182 per million is about 1.22 million Clef-flash tokens a day; divided by 21,818 it is about 458,000 Clef tokens a day. A typical customer message plus a three-question schema is 150 to 300 tokens. At 300 tokens, the free tier covers roughly 4,000 Clef-flash decisions a day, or 1,500 on Clef, every day, without a card on file. Images cost more, because an image is tokenized into the state; a 1024-pixel JPEG lands somewhere in the low thousands of tokens, so a photo-plus-text request is closer to 2,000 tokens and the free tier covers about 600 of those a day on Clef-flash.

Workload Tokens a month Clef-flash Clef Jev, for comparison
Jake's shop: 300 text messages a week at 300 tokens~390K$0 (inside the daily free tier)$0about $0.02
Jake's shop plus 150 photos a week at 2,000 tokens~1.7M$0 (about 57K tokens a day)$0Jev has no image input
A support desk: 5,000 tickets a day at 300 tokens45Mabout $0.75 (first 1.22M a day free)about $7.50about $1.89
A moderation pipeline: 1M items a day at 400 tokens12Babout $1,077about $2,877about $504

The table is the whole pricing argument in four rows. Below a few million tokens a month, Clef is free and Jev is not, so the per-token ratio is irrelevant. Above a few hundred million, Jev's lower rate wins back the money, and at a billion-item moderation scale the difference is real. Most readers of this page are in the first two rows. Ethan's version: "Free up to the point where you can afford it is the right shape for a price."

Two footnotes. The 10,000 Neurons are shared across every Workers AI model your account calls, so a chat model running on the same account eats the same allowance. And Cloudflare's own estimate of the free tier's reach for Clef, about 458K tokens, assumes you spend all of it on the big model; most people will not.

Your first Clef call: the exact request and the exact response

There are two doors. From inside a Cloudflare Worker, the AI binding is one line and needs no key, because the Worker already belongs to your account. From anywhere else, the REST endpoint takes your account ID in the path and an API token in the header. Both carry the same JSON body, and it is the Jev shape with one addition: the model field inside the body must match clef or clef-flash exactly, with the pattern enforced, so "model": "@cf/cloudflare/clef-flash" inside the body is rejected even though that is the model ID in the URL.

  1. Find your account ID. In the Cloudflare dashboard, open any zone or the Workers and Pages page; the account ID is in the right-hand column and in the URL. Copy it into a shell variable: export CLOUDFLARE_ACCOUNT_ID=....
  2. Make a Workers AI API key. Cloudflare calls it an API token: My Profile, API Tokens, Create Token, and give it the Workers AI Read permission. Copy it once; Cloudflare will not show it again. export CLOUDFLARE_AUTH_TOKEN=....
  3. Send the request. The example below is Cloudflare's own, with three questions of three types about one support message.
  4. Read result.answers. Workers AI wraps the model's reply in its standard envelope, so the answers are one level deeper than they are from Ollama or Jev.
curl https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/run/@cf/cloudflare/clef-flash \
  -X POST \
  -H "Authorization: Bearer $CLOUDFLARE_AUTH_TOKEN" \
  -d '{
    "model": "clef-flash",
    "state": "Checkout has been failing for every customer for the last hour.",
    "questions": {
      "urgent": { "type": "noul", "instructions": "Is this support request urgent?" },
      "team": {
        "type": "choice",
        "instructions": "Which team should handle this request?",
        "criteria": {
          "billing": "Payments, invoices, and refunds",
          "technical": "Outages, errors, and configuration",
          "sales": "Plans and upgrades"
        }
      },
      "severity": {
        "type": "score",
        "instructions": "How severe is the customer impact?",
        "criteria": ["No impact", "Minor", "Major", "Critical"]
      }
    }
  }'

The same call from a Worker, with the binding named AI in wrangler.toml:

const response = await env.AI.run("@cf/cloudflare/clef-flash", {
  model: "clef-flash",
  state: "Checkout has been failing for every customer for the last hour.",
  questions: {
    urgent: { type: "noul", instructions: "Is this support request urgent?" },
    team: {
      type: "choice",
      instructions: "Which team should handle this request?",
      criteria: {
        billing: "Payments, invoices, and refunds",
        technical: "Outages, errors, and configuration",
        sales: "Plans and upgrades"
      }
    },
    severity: {
      type: "score",
      instructions: "How severe is the customer impact?",
      criteria: ["No impact", "Minor", "Major", "Critical"]
    }
  }
});

What comes back is one answer per question, keyed by the names you chose. A noul answer carries a single probability that the statement is true. A choice answer carries the chosen key, a probability for every key, and a confidence figure that says how concentrated the distribution is. A score answer carries the chosen level, the probabilities across the scale, and a probability-weighted position. The usage block lists input tokens and an output count of zero. Through the REST door, all of that sits under result, beside success and errors; through the binding, you get the model's object directly.

Jake: "Where is the explanation of why it picked technical?"

Ethan: "There isn't one, and there will never be one. You get technical 0.97, billing 0.02, sales 0.01. That is the explanation. If a customer or an auditor needs a sentence, you write the sentence from the numbers and the criteria, or you hand the case to a chat model that can write. The decision model's job ends at the probabilities."

The three question types, and the limits that bite

Noul: a yes/no answered as a probability

One instruction, no criteria. "Is this support request urgent?" returns a probability between 0 and 1. Set your threshold from your own test set, not from 0.5; a value near 0.5 usually means your instruction is ambiguous, not that the model is confused.

Choice: one option from a named list

The criteria object maps option keys to descriptions. Clef accepts from 2 to 255 options per question, far more than the 26 Ollama exposes. The descriptions carry the meaning; the keys carry almost nothing, and the trap section of the Ollama guide shows what happens when you forget that.

Score: a position on an ordered scale

The criteria array lists 2 to 10 levels in order. The model returns the chosen level plus the full distribution, so you can use the expected value as a continuous score if you prefer. Scores are the weakest question type for every decision model on the index, Clef included; keep the scale short and the level descriptions concrete.

The limits are the kind you meet at 2 a.m. A request can hold 1 to 64 questions, each with an ID of letters, digits, underscores, dots or hyphens up to 100 characters. Images: at most four, each up to 4 MiB and 16 megapixels, 8 MiB decoded in total. The whole request body tops out at 13 MiB. The hosted context is 65,536 tokens; the model card's default for local use is 16,384 and is adjustable. And the body's model string must be exactly clef or clef-flash, surrounding whitespace allowed, nothing else.

Images: the thing Jev and the Ollama models cannot do

This is the feature that makes Clef more than a faster Jev. Both Clef sizes keep the vision encoder of their Qwen base, so the state can include photographs, screenshots, scanned forms or frames from a video, up to four per request, and the questions are asked about the pictures and the text together. Jev has no image input. Nimble and Tev1 on Ollama have no image input. Among hosted decision models on October 1, 2026, Clef is the one that can look.

For Jake that is the whole reason to add it. His booking form on Cloudflare Pages already accepts a photo. A Worker behind it can now ask Clef-flash four questions about the image before a human reads anything: damage as a choice between screen, battery, water, charging port, none visible; device as a choice between iPhone, Samsung, Pixel, other; readable as a noul, "is the damage clearly visible in this photo"; and urgency as a score. A screen job gets the screen quote template, a swollen battery gets the "do not charge it, bring it in today" reply, a photo of a cat gets routed to a human. None of that needed training. The criteria are the training.

Two launch-day gotchas worth knowing before you ship. Original phone photos are large: a developer testing on October 1 found that 1.3 MB PNGs failed with a token-overflow error, and resizing to 1024 pixels wide as JPEGs fixed it, so put a resize step in front of the model and keep the 4 MiB and 16-megapixel limits in mind. And image requests were slow on launch day, 13 to 30 seconds end to end, against under a quarter of a second for text; whether that was launch-day load or a steady cost of the vision path, design for it by asking the photo questions asynchronously rather than while a customer waits on a spinner.

Can you run Clef locally? The honest answer

Yes, in the sense that matters legally, and no, in the sense that matters on a Tuesday. The weights are Apache 2.0 and public. Cloudflare/clef-flash is 19.1 GB: four safetensors shards of the Qwen3.5-9B backbone, a 244 MB joint_head.safetensors, and a 23 KB Python file, joint_schema_model.py, that knows how to put them together. Cloudflare/clef is 55 GB with a 256 MB head. Cloudflare tested on a single H200 with torch 2.11 and transformers 5.10.2, and the two supported ways to run are the model card's Python (load_release_model(path, device="cuda")) and vllm serve "Cloudflare/clef". At full precision with a 64K context and one request at a time, Clef-flash needs about 41 GB of VRAM and Clef about 85 GB. That is a datacenter card or two, not a laptop.

Within four hours of launch the community had posted seven re-packagings of Clef-flash: GGUF files from two uploaders, MLX 4-bit and 8-bit for Apple Silicon, an FP8 and an EXL3. The largest GGUF set runs from a 3.54 GB IQ2_M to a 9.55 GB Q8_0 and a 17.9 GB bf16, with a 918 MB vision projector alongside. The Q4_K_M at 5.84 GB would fit any 8 GB graphics card and most 16 GB laptops, which is exactly the tier this series is written for. So why the asterisk?

Because those files are the Qwen backbone and its vision encoder, converted the way every chat model is converted. The 244 MB joint head, the part that turns hidden states into one logit per option, is not in them, and no GGUF runtime loads it. Mainline llama.cpp has no /v1/systemone endpoint yet; the pull request that adds one supports five other decision models (Laya, Julia-1, lev, OpenJev and Kev) and names Clef as a follow-up. Ollama's library has no clef tag. ollaya, the "Ollama for decision models" project, lists 27 models and Clef is not among them. The MLX repacks are in the same position until someone ports the head. You can download Clef-flash to your laptop tonight and run it as a chat model, which would be pointless because it was trained not to chat, but you cannot ask it a typed question. That will change, probably within weeks given how fast this category moves, and this page will get a dated line when it does.

Route Works on October 1, 2026? What you need
Workers AI, hostedYesA free Cloudflare account; nothing installed
transformers or vLLM with the official safetensorsYesAbout 41 GB VRAM (flash) or 85 GB (Clef) at bf16; Linux, CUDA, Python
GGUF in llama.cpp, LM Studio, JanLoads as a chat model; no decisionsThe joint head is not in the GGUF and no runtime reads it yet
Ollama /v1/systemoneNo tagOllama 0.35 serves Nimble and Tev1; Clef would need the head ported
MLX 4-bit or 8-bit on a MacBackbone only todaySame head problem; watch the repos
What to run instead on a laptopYesollama pull tev1 (4.5 GB) or nimble (9.5 GB); same request shape, text only, free

If you came here from a search for "run Clef locally" with a 16 GB laptop, the practical answer is the last row, and the RAM tiers in the Ollama guide tell you which one. If you are shopping for a machine specifically to run decision models at home, the honest laptop guide stands: nothing you can carry runs Clef-flash at full precision, and the quantized route is waiting on software, not hardware.

AI decision making examples: five things people build with Clef in the first week

  1. Routing in front of a big model. Every message hits Clef-flash first with two questions: "does this need a human" as a noul and "which model tier" as a choice between cheap, standard and expensive. The 90% that are simple go to a small chat model; the rest go to a frontier model. The decision costs a few hundred tokens at $0.09 per million; the savings on the other side are measured in dollars.
  2. Moderation that cannot be talked out of its answer. A chat model asked "is this harassment" can be argued with inside the message itself. A decision model scores the options and has no channel for the message to negotiate through. Clef's confidence figure lets you send the uncertain middle to a person.
  3. Domain and URL classification. Cloudflare's own Threat Intelligence team uses Clef to classify website domains, and quotes 2.2 seconds for a classification that took 4.7 seconds with a 120B open chat model. If you run a mail filter, a kids' network or an ad-fraud check, this is the same job.
  4. Photo triage. Jake's case: damage type, device, legibility, urgency, from the picture, before a human reads the message. Insurance claims, returns departments and property managers have the identical shape.
  5. Invoice and form fields as choices. "Which cost center," "which tax category," "is this a duplicate of an invoice we have seen" are choice and noul questions over a scanned document. Clef's 64.7% on the invoice-processing workflow set says it is not an accountant; it is a fast first pass that a person corrects, which is still most of the work.

What ties the five together is the question "how does AI make a decision" in its plainest form. A decision model does not reason toward an answer; it was trained so that, given a state and a set of options, the option it has learned belongs to that state gets the highest score. Automated decision making built on that is auditable in a way chat-based routing never was, because every decision leaves a complete probability record and the criteria you wrote are the policy. Explaining decisions made with AI, in the regulatory sense, becomes "here are the options we offered, here is the probability it assigned each, here is the threshold we set," which is a page a lawyer can read.

The trap carries over: option names still steer the answer

The Ollama guide ends on a finding from a September study of Jev-style models: changing option keys from 0/1 to no/yes while leaving the descriptions identical flipped about 70 of every 100 answers, with a type-error rate of exactly zero. Type-safe is not error-free. Clef is trained better than the models in that study, calibrated with a Brier loss and a reinforcement stage built for exactly this, and its 97% on CLINC150 with out-of-scope detection says it handles "none of these" better than most. None of that repeals the rule. Keys that lean (good/bad, urgent/whatever) will lean the answer. Descriptions that are vague will produce probabilities near the middle. A schema that worked on Tev1 will produce different probabilities on Clef, so your thresholds move when your model does.

The habit that makes any of these models safe is the same one the last guide asked for. Keep 50 to 100 real items with the label you would have given. Run them through the new model before you point production at it. Count. Set the threshold where your own numbers say, and re-run the fifty when Cloudflare ships a new version, because the model ID will not change and the probabilities will. Ethan: "A model you never counted is a rumor. A rumor with a probability attached is still a rumor."

What Cloudflare keeps, and what you can turn on

Cloudflare's statement for Clef is unusually direct: they do not read, store or train on your requests or responses unless you opt into the fine-tuning program. Requests are processed on Cloudflare's GPUs in the data center nearest the request, which is the "edge" in the launch language and the reason the latency figures are what they are. If you want a record, AI Gateway sits in front of Workers AI and can log every request and response to your own account, with caching and rate limits, and it is off until you turn it on. Compared with the local route, where nothing leaves the machine, this is a trade: your customers' messages and photos do travel to a third party. For many small businesses whose website already lives on Cloudflare, that third party is already holding the data, and the trade is small. For a clinic or a law office, the local clerk and the hosted clerk are different legal animals, and the decision belongs to whoever signs your privacy policy, not to a benchmark.

The RL fine-tuning platform: what is real today

Alongside Clef, Cloudflare announced a reinforcement-learning fine-tuning platform for decision models, and the announcement is honest about its stage. Today it is a design-partner program with Cloudflare's forward-deployed engineers doing the work; a self-serve version is planned, with no date. The pipeline they describe uses the pieces they already sell: AI Gateway captures your real requests and responses, Workers AI generates rollouts, Cloudflare Containers act as a sandbox that scores them, a new Trainer component updates the weights, and the result deploys back onto Workers AI through Bring Your Own Model. The promise is a Clef tuned on your own decisions, with your own data never leaving your account's boundary.

For a small business this is a feature to watch, not to wait for. Together trained Tev1 for about $17 and 25 minutes on a public recipe, and the open-weight models on the index were built by individuals in their spare time. If the self-serve platform lands at a price in that range, it will matter. Until there is a price, the practical path is the one already in your hands: write better criteria, test on fifty, and move the questions Clef-flash gets wrong up to Clef.

Decision models in October 2026: who is shipping what

 Three weeks that built a category

  • Mid-September: TypeSafe announces Jev, a hosted System One API at $0.042 per million input tokens. Within two weeks there are more than twenty open-weight imitations on Hugging Face and a public leaderboard.
  • September 28: Ollama 0.35.0 adds /v1/systemone and ships Nimble 9B and Tev1 4B and 0.8B. Nimble passes 11,000 pulls in its first days.
  • September 29: OpenAI announces a hosted Decisions API at DevDay, running on a GPT-6 Luna variant in about 150 milliseconds, limited preview, no price.
  • October 1: Cloudflare ships Clef and Clef-flash on Workers AI with open weights, the first decision models with image input and a free tier. The same day, llama.cpp's own organization posts GGUF conversions of Kev 4B, Laya, lev and Julia-1 for the /v1/systemone server that is still in review, and the community posts seven repacks of Clef-flash within four hours.
  • Still to come: Clef on the independent leaderboard; a GGUF or MLX runtime that loads the joint head; an ollama pull clef; a price on OpenAI's API; a date on Cloudflare's self-serve fine-tuning.

If you are deciding where to place a bet for a product, the shape of the market is now clear. The request format is settled; everyone speaks Jev's. The hosted tier has three players with three prices: TypeSafe cheapest per token, Cloudflare free at the bottom and fastest, OpenAI not yet priced. The local tier is Ollama today and llama.cpp soon, with the open models two to three sizes smaller than the hosted ones and scoring accordingly. Writing your code against the shared request shape means you can move between all of them by changing a URL, which is a luxury the chat-model world never offered.

When something goes wrong: the errors you will actually meet

  1. A 400 with a complaint about model. The body's model string must be clef or clef-flash. People put the full @cf/cloudflare/clef-flash ID in the body because it is in the URL; the URL wants the long form, the body wants the short one.
  2. A question ID rejected. IDs are letters, digits, underscore, dot and hyphen, up to 100 characters. Spaces and slashes fail. Fix the key, not the model.
  3. Token overflow on an image. Resize to about 1024 pixels on the long side and send JPEG. Keep each image under 4 MiB and 16 megapixels, the set under 8 MiB decoded, the whole body under 13 MiB.
  4. "Neurons exhausted" or a 429 after a busy day. You have used the 10,000 free Neurons, which every Workers AI model on the account shares. Either add the Workers Paid plan, where the overage is $0.011 per 1,000 Neurons, or wait for the daily reset.
  5. Probabilities near 0.5 on a noul, or a flat distribution on a choice. Not an error. Your instruction or your descriptions are ambiguous, or the item is genuinely borderline. Tighten the wording, add an explicit "none of these" option, and route low-confidence items to a person.
  6. Different answers from Clef and Clef-flash on the same item. Expected. They are different models. Pick one per question using your 50-item test and do not mix them for the same question without re-testing the threshold.
  7. A slow image request. Launch-day image calls ran 13 to 30 seconds. Ask photo questions in the background and show the customer a receipt, not a spinner.

Update, October 2, 2026: llama.cpp merged its own /v1/systemone endpoint today, with ready-made GGUFs for five open decision models: OpenJev (27B, reads screenshots, non-commercial license), Kev-4B and Lev (about 3 GB, Apache 2.0), Laya and Julia-1 (under 1 GB). Sizes, the license trap and the exact commands: OpenJev locally: llama.cpp now runs Jev-style decision models.

Cloudflare Clef: the questions people are typing this week

What is Cloudflare Clef?

Clef is a family of two decision models trained by Cloudflare's Workers AI team and released October 1, 2026: Clef, 27 billion parameters on a Qwen3.8-27B base, and Clef-flash, 9 billion on Qwen3.5-9B. Given text and up to four images plus typed questions, each returns a choice, a yes/no probability or a score with a probability for every option, in one pass, without writing any text. Both run hosted on Workers AI and are Apache 2.0 with weights on Hugging Face.

What is decision AI, or a decision model?

A decision model is a neural network trained to pick from options you supply rather than to generate text. You send a state (text, JSON or images) and questions with their allowed answers; it scores every allowed answer and returns the distribution. It cannot invent an option or explain itself. The category was named System One by TypeSafe when it launched Jev in September 2026, and Clef, Nimble, Tev1 and Laya all follow the same request shape.

Is Cloudflare Clef free?

Up to a point that covers most small uses. Every Workers AI account gets 10,000 Neurons a day at no charge, which is about 1.22 million Clef-flash input tokens or 458,000 Clef tokens, roughly 4,000 short text decisions a day on Clef-flash. Beyond that, Clef-flash is $0.09 and Clef $0.24 per million input tokens on the Workers Paid plan, and output tokens are never billed.

Clef vs Jev: which is better?

Clef-flash is about 13 times faster and Clef 2.5 times faster on Cloudflare's measurements, both can read images, and both have a free tier; Jev is 2 to 6 times cheaper per token and stronger on reasoning-heavy questions (GPQA Diamond 78.3 against Clef's 48.0). Cloudflare reports Clef at 61.2 on the Decision Index against Jev's 57.9, but that run is Cloudflare's own and not yet on the public leaderboard. For bounded classification at small to medium volume, Clef-flash is the better bet today; for very large volumes, Jev's price wins.

Is Jev AI free, and what does TypeSafe charge?

Jev is a paid, hosted API in early access at $0.042 per million input tokens with output free; no free tier has been published and there are no downloadable weights. Clef's free allocation and the Ollama models are the free routes to the same request shape.

Can I run Clef locally?

The weights are open and you can, with a large GPU: about 41 GB of VRAM for Clef-flash and 85 GB for Clef at full precision, using Cloudflare's Python loader or vLLM. On a laptop, not yet. The GGUF and MLX repacks posted on launch day contain the Qwen backbone but not the 244 MB joint head that makes the decision, and no desktop runtime loads it. Run Tev1 or Nimble on Ollama for local decisions and use Clef hosted for images.

Does Clef work with Ollama?

Not as of October 1, 2026. Ollama's library has no clef tag and its /v1/systemone endpoint serves Nimble and Tev1. The request shape is identical, so code written for Ollama's endpoint can be pointed at Workers AI by changing the URL, adding the Authorization header and reading answers from inside the result envelope.

What is Clef-flash, and should I use it instead of Clef?

Clef-flash is the 9-billion-parameter model, a third of the price and about five times faster than Clef, and it scores higher on several of Cloudflare's headline benchmarks including BFCL (98.76% against 98.47%). Start with Clef-flash for short, well-defined questions and move a question to Clef only when your own test set shows the small model getting it wrong, typically on long states or many options.

How does AI make a decision in Clef?

The backbone reads the state and every question; a joint schema head then emits one logit per allowed option for each question, and a softmax turns those into probabilities. There is no token-by-token generation and no reasoning chain. Training with label smoothing, a Brier loss and a reinforcement-learning stage for calibration is meant to make the probabilities honest, so a 0.9 is right about nine times in ten on data like the training set.

Can Clef explain its decision?

No. The output is the chosen option, the probability for every option and a confidence value. If you need a written reason, generate it separately with a chat model from the probabilities and your criteria, or treat the criteria plus the distribution as the explanation; for audit purposes that record is more complete than a paragraph.

Does Clef read images?

Yes, both sizes. A request can include up to four images, each up to 4 MiB and 16 megapixels, 8 MiB decoded in total, and the questions are answered about the images and text together. Jev and the Ollama decision models cannot do this. Resize phone photos to about 1024 pixels and send JPEG to avoid token-overflow errors.

Is Cloudflare Workers AI free, and what are the free tier limits for Clef?

Workers AI is free up to 10,000 Neurons a day on every account, with no card required, and Clef is 21,818 Neurons per million input tokens and Clef-flash 8,182. That daily allowance is shared across all Workers AI models you call; past that, the Workers Paid plan charges $0.011 per 1,000 Neurons, which works out to $0.24 and $0.09 per million tokens respectively.

Is Clef open source?

The weights and the loader code are Apache 2.0 on Hugging Face, which is open weights in the usual sense. The training data is not published, which Cloudflare has confirmed. Call it open-weight if the distinction matters to your compliance team.

What is System One in AI?

A term borrowed from Daniel Kahneman's Thinking, Fast and Slow: System 1 is fast, intuitive judgment and System 2 is slow, deliberate reasoning. TypeSafe used it to describe Jev, a model that returns a typed decision in a single pass, and Ollama named its endpoint /v1/systemone after it. Clef is a System One model; a chat model with a long thinking chain is System Two.

What is the Jev Decision Index?

The public leaderboard for decision models, a frozen suite of 120,340 requests across 43 benchmarks in five equal-weight areas, run with fixed prompts and no per-model tuning. Its headline number is a chance-normalized skill score, so random guessing scores near zero. The public results file from September 28, 2026 has Jev at 57.91, Nimble at 39.57 and Tev1 4B at 29.24; Cloudflare's own run of the suite puts Clef at 61.2 and Clef-flash at 57.1.

Which is best for text classification: Clef, Jev or a local model?

For bounded labels at small to medium volume, Clef-flash: fast, free under the daily allowance, and strong on intent and routing benchmarks. For very large volumes where tokens are the cost, Jev. For data that cannot leave the building, Tev1 or Nimble on Ollama, accepting lower scores. For long documents or labels that need reasoning, a chat model with structured output still wins on accuracy, at a higher price per item.

If you are reading this because an inbox or a photo queue has been quietly eating your evenings, the news is good and uncomplicated. The clerk that reads pictures now exists, it is free at the scale most of us work at, and it speaks the same four lines of JSON as the one already on your PC. Open a Worker, paste the request above, change the questions to yours, and count fifty. Jake's booking form now sorts the photographs before he has finished his first coffee, and the swollen batteries get the urgent reply first instead of last. If a number here has moved by the time you read it, tell me; pages like this stay right because readers write in.

 If you keep one line from this page

Clef is free until you can afford it, reads the photos nothing else in this category can, and is "open source" in a way that does not yet run on anything you own; start with Clef-flash, keep the local clerk, and count fifty before you trust a threshold.

Open weights are a promise. A runtime is a fact. Today the promise is on Hugging Face and the fact is on Workers AI.

Revision note. Written October 1, 2026, the day Clef shipped, with prices, limits and scores as published that day. If you are reading this with a photo queue and a free Cloudflare account, the first call is ten minutes away and costs nothing.

Related