· 6 min read
Cloudflare Clef: decision models for AI agents on Workers AI
Cloudflare Clef is an open-weight decision model that returns calibrated probabilities instead of text. How it works and when to use it over an LLM.

Cloudflare Clef is a pair of open-weight decision models, released on Workers AI on 1 October 2026, that answer typed questions about an input instead of generating text. You send a piece of state and a schema of questions, and Clef returns a probability for each possible answer. It is built for the small, repeated choices inside an agent: is this urgent, which team owns it, how bad is it.
That sounds like classification, and it is. What is new is the packaging: a schema-driven API, calibrated probabilities, image input and open weights, at latencies low enough to sit in the request path. Here is how it works and where it fits.
What is a decision model?
A decision model is a model that picks from a bounded set of answers and reports how confident it is, rather than writing free-form output. Cloudflare credits Typesafe AI's Jev System One model with introducing the format, and Clef is designed to be API-compatible with it, so code written for one can switch to the other.
The difference from a chat model is in the shape of the output. Ask a general LLM "is this ticket urgent?" and you get text back. You then parse it, hope it said "yes" or "no" and not "it depends", and you have no real sense of how sure it was. A decision model can only return one of the answers you defined, and every answer comes with a probability. Your code branches on a number, not on a string.
How does Cloudflare Clef work?
Clef comes in two sizes:
| Model | Base | Workers AI ID | Notes |
|---|---|---|---|
| Clef | Qwen 3.8, 27B | @cf/cloudflare/clef | Higher accuracy, about 209 ms median latency |
| Clef-flash | Qwen 3.5, 9B | @cf/cloudflare/clef-flash | Built for latency-critical paths, about 39 ms median |
The latency figures are Cloudflare's own measurements. Both models have a 65,536-token context window, accept up to four images per request (PNG, JPEG or WebP), and the weights are published on Hugging Face under the Apache 2.0 license.
The interesting part is how the answer is produced. A normal LLM generates output one token at a time, which is where most of its latency comes from. Clef does a single prefill pass, meaning the model reads the whole input once and builds its internal representation, and then scores every valid answer in your schema in parallel. There is no generation step and no intermediate text. That is why a 9B model can answer in tens of milliseconds, and why the output can never fall outside the schema.
Calibration is the other half. A calibrated model is one whose probabilities match reality: across many answers marked 80%, roughly 80% should be correct. Cloudflare trained Clef with a loss that penalises overconfident probabilities, followed by a reinforcement learning stage aimed at calibration. If the numbers are trustworthy, you can set thresholds on them.
What does a Clef request look like?
A request has three parts: the model selector, the state you want judged, and a questions object with between 1 and 64 named questions. This is the example from Cloudflare's announcement:
curl https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/run/@cf/cloudflare/clef \
-X POST \
-H "Authorization: Bearer $CLOUDFLARE_AUTH_TOKEN" \
-d '{
"model": "clef",
"state": "Checkout has been failing for every customer for the last hour.",
"questions": {
"urgent": { "type": "noul", "instructions": "Is this support request urgent?" },
"team": {
"type": "choice",
"instructions": "Which team should handle this request?",
"criteria": {
"billing": "Payments, invoices, and refunds",
"technical": "Outages, errors, and configuration",
"sales": "Plans and upgrades"
}
},
"severity": {
"type": "score",
"instructions": "How severe is the customer impact?",
"criteria": ["No impact", "Minor", "Major", "Critical"]
}
}
}'
There are three question types:
- noul is a yes/no question. The answer is the probability that the answer is yes.
- choice picks one option from a named set, and returns the chosen option with a probability for each.
- score places the state on an ordered scale, such as the four severity levels above, and returns a probability-weighted score.
The response contains an answers object keyed by your question IDs, plus a usage object. The state does not have to be a string; the docs allow structured data such as objects and arrays, which is handy when the thing being judged is an event or a database row.
Inside a Worker, the same call goes through the AI binding:
const response = await env.AI.run("@cf/cloudflare/clef", {
model: "clef-flash",
state: ticket.body,
questions: {
urgent: { type: "noul", instructions: "Is this support request urgent?" },
},
});
console.log(response.answers.urgent);
Pricing on Workers AI is listed at $0.24 per million input tokens for Clef. Because nothing is generated, output tokens are not the cost driver they are with a chat model.
When should you use a decision model instead of an LLM?
Use one when the set of possible answers is known in advance and you call it often. Common cases:
- Routing. Sending a ticket, email or agent step to the right handler.
- Gating. Deciding whether a request needs a human, a stronger model, or no action.
- Moderation and triage. Labelling submissions, including ones with screenshots.
- Agent tool choice. Picking the next action from a fixed list before handing off to a larger model.
The probabilities are what make this practical. Instead of trusting every answer, you can act automatically above a threshold and escalate below it. For example, auto-route a ticket when the top team has a probability above 0.9, and send everything else to a person or to a bigger model. You get cheap, fast handling for the easy majority and spend effort only on the uncertain cases.
Stick with a general LLM when the answer is open-ended, when you need an explanation, or when the options change on every call in ways you cannot express as a schema. A decision model will not summarise, draft or reason in text. It also only knows about the options you give it, so a badly written criteria description will produce confident answers to the wrong question.
A useful pattern is to put both in the same pipeline: Clef-flash decides what kind of request this is, and only the requests that need real generation reach a chat model. That keeps your expensive tokens for the work that needs them, the same instinct behind reusing cached prompt prefixes to cut input costs.
What about fine-tuning Clef?
Alongside the models, Cloudflare announced a reinforcement learning fine-tuning platform. It captures real requests and responses through AI Gateway, generates rollouts on Workers AI, scores them in a sandbox running on Cloudflare Containers, trains updated weights, and redeploys the result on Workers AI. For now it runs as a hands-on program with Cloudflare's engineers, and self-serve access is planned for later. Because the weights are open, you can also fine-tune or host Clef yourself.
Key takeaways
- Cloudflare Clef answers typed questions with probabilities instead of generating text.
- It scores all valid answers in one prefill pass, so output always fits your schema and latency stays low.
- Clef (27B) and Clef-flash (9B) run on Workers AI and ship as Apache 2.0 open weights.
- Use the probabilities as thresholds: automate confident decisions, escalate uncertain ones.
- Keep general LLMs for open-ended work, and use a decision model in front of them for routing.
FAQ
Is Cloudflare Clef open source?
Yes. Both Clef and Clef-flash are released with open weights on Hugging Face under the Apache 2.0 license, and they are also available as hosted models on Workers AI.
Can Clef read images?
Yes. A request can include up to four images in PNG, JPEG or WebP format, which lets you make decisions about screenshots, photos or documents alongside text.
How many questions can one Clef request answer?
Between 1 and 64. All of them are answered from the same pass over the input, so asking several related questions in one request is cheaper than making separate calls.