An AI decision model is a model that answers a fixed set of typed questions with probabilities your code can act on. Unlike an LLM asked to "act as a classifier," it does not write free-form text you then parse. Cloudflare's Clef (27B) and Clef-flash (9B), announced October 1, 2026, are the concrete example: Apache 2.0 weights, hosted on Workers AI, and built for hot-path choices like ticket routing, urgency, and tool choice.
Last verified: October 5, 2026. Specs below (sizes, Apache 2.0 weights, Workers AI hosting, System One / Jev-compatible API shape, RL fine-tune design-partner path) come from Cloudflare's launch blog and Workers AI changelog. Latency and head-to-head scores vs other decision models are Cloudflare's published numbers from that launch week. We did not re-run their benches. Product pricing and third-party Jev claims beyond Cloudflare's own API-compatibility note are omitted. If a comparison is older than about 90 days after this launch, distrust the vendor table.
What kinds of models sit in an agent loop?
| Category | What it returns | Best for |
|---|---|---|
| Chat / reasoning LLM | Free-form text, plans, tool-call prose | Open-ended work: write, explain, invent steps |
| Classic classifier | Labels you trained for | Stable taxonomies you retrain when labels change |
| AI decision model | Typed answers with probabilities over options you define | Hot-path decide/act: route, escalate, pick a tool, score severity |
Most agent demos put an LLM in every slot. That works for demos. It gets expensive and flaky when the job is "is this urgent, and which team owns it?" on every ticket.
Decision model vs LLM-as-classifier
An AI decision model takes a state (text, JSON, sometimes images) plus a schema of questions. Each question is typed. Cloudflare's System One shape (the family Clef follows) uses three question types: yes/no (noul), pick-one (choice), and ordered score (score). The model returns a probability for every allowed answer. Your code branches on those numbers. There is no prose to scrape and no "reasoning tokens" to wait on for the decision itself.
An LLM forced to act as a classifier is different. You prompt it to output JSON labels. You still get free-form generation underneath. Formats drift. Confidence is mostly theater unless you calibrate it yourself. Latency and cost scale with how much text it invents before the label.
Decision rules:
| Pick this | When |
|---|---|
| Decision model | The answers are a closed set you already know; you need probabilities for thresholds; the call sits on the hot path (routing, urgency, tool choice, guardrails) |
| LLM | The next step needs new text, multi-step reasoning, or a tool plan you cannot schema in advance |
| Neither alone | You need both: decide with a decision model, then hand off to an LLM to write the reply or call tools |
That split is the whole point of this page. Clef is only the worked example.
How a decision model runs one unit of work
Walk one support message through the loop Cloudflare documents:
- State in. "Checkout has been failing for every customer for the last hour."
- Questions in. Is it urgent? (
noul) Which team? (choice: billing / technical / sales) How severe? (scoreover an ordered rubric) - Answers out. Probabilities per option, keyed by the same question ids. Your agent reads
urgent,team, andseverityas structured fields. - Act or escalate. Route the ticket, block a request, call a tool, or defer to a human when confidence is low.
You should now see why this is not "chat with a system prompt." The contract is state + schema in, typed probabilities out.
For the wider agent stack around that decide step (tools and memory), see what an MCP server is and how AI agent memory works. More tool explainers sit in the AI tools hub.
Cloudflare Clef as a concrete example
On October 1, 2026 Cloudflare released Clef and Clef-flash, the first models trained by the Workers AI team and hosted on Workers AI as @cf/cloudflare/clef and @cf/cloudflare/clef-flash. Weights are open under Apache 2.0 on Hugging Face. Cloudflare describes them as Jev-API compatible (System One API): you can point an existing Jev-shaped integration at Clef by changing endpoint and model.
| Model | Size | Cloudflare says it is best for | Context window |
|---|---|---|---|
| Clef | 27B | Highest-precision decisions | 64K tokens (docs: 65,536) |
| Clef-flash | 9B | Latency-critical, hot-path decisions | 64K tokens (docs: 65,536) |
Verified extras from Cloudflare primary sources only:
- Vision: optional images (up to four) alongside the state; Cloudflare positions this as a Clef extension vs text-only decision models they compare against.
- Batching: up to 64 questions per request.
- RL fine-tune: Cloudflare is opening a reinforcement learning fine-tuning path starting with design partners / FDEs, not a self-serve click-to-tune product as of the launch posts.
- Latency (Cloudflare's numbers): on their 43 benchmark runs they report median latency of 209.3 ms for Clef and 38.8 ms for Clef-flash. Treat that as vendor-published, not independently rechecked here.
What doesn't get said. Clef is not a general chat model. It will not write your customer email. It will not invent a new workflow. If your schema is wrong, the probabilities are still "correct" against a bad question. And RL fine-tuning at launch is a design-partner motion: you should not plan your roadmap as if self-serve custom Clef weights ship tomorrow.
Verdict: use Clef (or Clef-flash) when you already know the decision options and need fast, typed probabilities on Workers AI or local Apache 2.0 weights. Skip it when the agent still needs to generate the plan, not just pick among plans.
When to put a decision model in the hot path
| Situation | Use a decision model? | Why |
|---|---|---|
| Ticket routing and urgency | Yes | Closed teams and urgency flags; thresholds beat prompt theater |
| Tool / skill choice before an LLM acts | Yes | Guardrail check in tens of milliseconds (Cloudflare's framing) |
| Trust and safety score against your rubric | Yes | Ordered score questions map cleanly |
| Drafting the reply or writing code | No | That is LLM work after the decide step |
| Brand-new taxonomy you cannot list yet | Not yet | Schema the options first, or stay with an LLM until the set stabilizes |
| One-off research question | No | Open-ended; decision models waste the constraint |
Cloudflare's own internal example is threat-intel style domain classification with Browser Run: they report Clef fetching, rendering, and classifying a domain in 2.2 seconds in that workflow versus 4.7 seconds for their general LLM gpt-oss-120b in the same setup. That is a Cloudflare anecdote from the launch blog, not a universal speed law.
Which one should I pick?
| Your situation | Pick | Why |
|---|---|---|
| You need structured route / escalate / score on every request | Decision model (e.g. Clef or Clef-flash) | Typed probabilities; no parse step |
| You need the lowest latency Cloudflare lists for this family | Clef-flash (9B) | Cloudflare positions it for hot-path latency |
| You need max precision in that same family | Clef (27B) | Cloudflare positions it for highest-precision decisions |
| You need open-ended reasoning or generated text | LLM | Decision models do not write |
| You need both decide and act | Decision model then LLM | Decide on the hot path; generate after |
| You are still learning agent architecture | Learn the decide/act split first | Tools and memory matter as much as the model brand |
If you are comparing learning paths for shipping agents (not picking a Cloudflare SKU), an advisor can help you see which program fits: AI Engineering for career changers and semi-technical pros who want to build production agents, not only prompt chatbots.
