4Geeks chosen to deliver AI education in the Bahamas alongside Harvard, Oxford, and Columbia.See more
7 min read

AI Decision Model: What It Is and When Agents Should Use One

An AI decision model returns typed probabilities for routing and urgency, not free-form text. When to use one vs an LLM, with Clef as the example.

An AI decision model is a model that answers a fixed set of typed questions with probabilities your code can act on. Unlike an LLM asked to "act as a classifier," it does not write free-form text you then parse. Cloudflare's Clef (27B) and Clef-flash (9B), announced October 1, 2026, are the concrete example: Apache 2.0 weights, hosted on Workers AI, and built for hot-path choices like ticket routing, urgency, and tool choice.

Last verified: October 5, 2026. Specs below (sizes, Apache 2.0 weights, Workers AI hosting, System One / Jev-compatible API shape, RL fine-tune design-partner path) come from Cloudflare's launch blog and Workers AI changelog. Latency and head-to-head scores vs other decision models are Cloudflare's published numbers from that launch week. We did not re-run their benches. Product pricing and third-party Jev claims beyond Cloudflare's own API-compatibility note are omitted. If a comparison is older than about 90 days after this launch, distrust the vendor table.


What kinds of models sit in an agent loop?

CategoryWhat it returnsBest for
Chat / reasoning LLMFree-form text, plans, tool-call proseOpen-ended work: write, explain, invent steps
Classic classifierLabels you trained forStable taxonomies you retrain when labels change
AI decision modelTyped answers with probabilities over options you defineHot-path decide/act: route, escalate, pick a tool, score severity

Most agent demos put an LLM in every slot. That works for demos. It gets expensive and flaky when the job is "is this urgent, and which team owns it?" on every ticket.


Decision model vs LLM-as-classifier

An AI decision model takes a state (text, JSON, sometimes images) plus a schema of questions. Each question is typed. Cloudflare's System One shape (the family Clef follows) uses three question types: yes/no (noul), pick-one (choice), and ordered score (score). The model returns a probability for every allowed answer. Your code branches on those numbers. There is no prose to scrape and no "reasoning tokens" to wait on for the decision itself.

An LLM forced to act as a classifier is different. You prompt it to output JSON labels. You still get free-form generation underneath. Formats drift. Confidence is mostly theater unless you calibrate it yourself. Latency and cost scale with how much text it invents before the label.

Decision rules:

Pick thisWhen
Decision modelThe answers are a closed set you already know; you need probabilities for thresholds; the call sits on the hot path (routing, urgency, tool choice, guardrails)
LLMThe next step needs new text, multi-step reasoning, or a tool plan you cannot schema in advance
Neither aloneYou need both: decide with a decision model, then hand off to an LLM to write the reply or call tools

That split is the whole point of this page. Clef is only the worked example.


How a decision model runs one unit of work

Walk one support message through the loop Cloudflare documents:

  1. State in. "Checkout has been failing for every customer for the last hour."
  2. Questions in. Is it urgent? (noul) Which team? (choice: billing / technical / sales) How severe? (score over an ordered rubric)
  3. Answers out. Probabilities per option, keyed by the same question ids. Your agent reads urgent, team, and severity as structured fields.
  4. Act or escalate. Route the ticket, block a request, call a tool, or defer to a human when confidence is low.

You should now see why this is not "chat with a system prompt." The contract is state + schema in, typed probabilities out.

For the wider agent stack around that decide step (tools and memory), see what an MCP server is and how AI agent memory works. More tool explainers sit in the AI tools hub.


Cloudflare Clef as a concrete example

On October 1, 2026 Cloudflare released Clef and Clef-flash, the first models trained by the Workers AI team and hosted on Workers AI as @cf/cloudflare/clef and @cf/cloudflare/clef-flash. Weights are open under Apache 2.0 on Hugging Face. Cloudflare describes them as Jev-API compatible (System One API): you can point an existing Jev-shaped integration at Clef by changing endpoint and model.

ModelSizeCloudflare says it is best forContext window
Clef27BHighest-precision decisions64K tokens (docs: 65,536)
Clef-flash9BLatency-critical, hot-path decisions64K tokens (docs: 65,536)

Verified extras from Cloudflare primary sources only:

  • Vision: optional images (up to four) alongside the state; Cloudflare positions this as a Clef extension vs text-only decision models they compare against.
  • Batching: up to 64 questions per request.
  • RL fine-tune: Cloudflare is opening a reinforcement learning fine-tuning path starting with design partners / FDEs, not a self-serve click-to-tune product as of the launch posts.
  • Latency (Cloudflare's numbers): on their 43 benchmark runs they report median latency of 209.3 ms for Clef and 38.8 ms for Clef-flash. Treat that as vendor-published, not independently rechecked here.

What doesn't get said. Clef is not a general chat model. It will not write your customer email. It will not invent a new workflow. If your schema is wrong, the probabilities are still "correct" against a bad question. And RL fine-tuning at launch is a design-partner motion: you should not plan your roadmap as if self-serve custom Clef weights ship tomorrow.

Verdict: use Clef (or Clef-flash) when you already know the decision options and need fast, typed probabilities on Workers AI or local Apache 2.0 weights. Skip it when the agent still needs to generate the plan, not just pick among plans.


When to put a decision model in the hot path

SituationUse a decision model?Why
Ticket routing and urgencyYesClosed teams and urgency flags; thresholds beat prompt theater
Tool / skill choice before an LLM actsYesGuardrail check in tens of milliseconds (Cloudflare's framing)
Trust and safety score against your rubricYesOrdered score questions map cleanly
Drafting the reply or writing codeNoThat is LLM work after the decide step
Brand-new taxonomy you cannot list yetNot yetSchema the options first, or stay with an LLM until the set stabilizes
One-off research questionNoOpen-ended; decision models waste the constraint

Cloudflare's own internal example is threat-intel style domain classification with Browser Run: they report Clef fetching, rendering, and classifying a domain in 2.2 seconds in that workflow versus 4.7 seconds for their general LLM gpt-oss-120b in the same setup. That is a Cloudflare anecdote from the launch blog, not a universal speed law.


Which one should I pick?

Your situationPickWhy
You need structured route / escalate / score on every requestDecision model (e.g. Clef or Clef-flash)Typed probabilities; no parse step
You need the lowest latency Cloudflare lists for this familyClef-flash (9B)Cloudflare positions it for hot-path latency
You need max precision in that same familyClef (27B)Cloudflare positions it for highest-precision decisions
You need open-ended reasoning or generated textLLMDecision models do not write
You need both decide and actDecision model then LLMDecide on the hot path; generate after
You are still learning agent architectureLearn the decide/act split firstTools and memory matter as much as the model brand

If you are comparing learning paths for shipping agents (not picking a Cloudflare SKU), an advisor can help you see which program fits: AI Engineering for career changers and semi-technical pros who want to build production agents, not only prompt chatbots.

Building agents, or only prompting chatbots?

Tell us where you are today and an advisor will help you see which program fits: AI Engineering for shipping agents, RAG, and production AI apps.

Frequently Asked Questions