Skip to content

Decision Models: When Scoring Beats Generating

#decision-models #llama.cpp #cloudflare-clef #systemone-api #llm-inference #agent-observability

Every LLM pipeline that makes a decision has the same ugly middle step: you prompt, you get a paragraph, you parse it, you retry when the format breaks. Routing a support ticket. Classifying a document. Deciding whether an agent's last step was safe. For these jobs you're buying hundreds of output tokens to get one label, and the label arrives wrapped in prose.

A new class of models removes that step. Decision models don't generate text. You send a state, a list of typed questions, and the allowed options; they read the input once and return a probability for every option. llama.cpp has standardized this at the /v1/systemone endpoint, and Cloudflare just released Clef, a 27B multimodal model built for exactly this API.

Key Numbers

  • 0 output tokens per answer. The API reports output_tokens: 0 because there is no generation to count.
  • 5 models ship in llama.cpp's decision collection, from 144M to 27B. Cloudflare's Clef makes six in the System One family.
  • 3 ms is Julia-1's median time to answer one question on an NVIDIA RTX PRO 6000. Clef sits at 209 ms but adds image and video input.
  • Apache-2.0 on every model here except OpenJev, which is CC BY-NC 4.0 and not free for commercial use.

Why parsing generated text is the wrong bottleneck ​

A chat model spends one forward pass per output token. Ask it to justify a yes/no decision with a 200-token explanation and you've paid 200 passes for one bit of information. A decision model reads the state once, scores your options, and stops. No sampling, no max_tokens guessing, no JSON repair loop.

The confidence story matters just as much. A chat model's "I think this is billing" is a vibe you have to trust. A decision model's billing: 0.90 is a calibrated probability from a softmax over exactly the options you defined. That single difference turns a prompt into a threshold you can act on: route the confident answers, escalate the rest.

The answer is always one of your options. That constraint sounds limiting. For routing, content moderation, safety checks, and agent control, it's the whole point. A decision model can't drift off-script, because going off-script isn't representable in its output.

How a decision model works ​

Under the hood these are still transformers. The difference is the head. Clef, for instance, is Qwen3.8-27B post-trained with a joint schema head, a small transformer that reads the backbone's final hidden states, routes evidence from the state to each question, and emits one logit per allowed option. A softmax over each question's options turns those logits into probabilities.

The shared contract is a state plus typed questions. Three question types cover most real decisions:

TypeYou sendYou get
choicenamed options, with optional descriptionstop option plus a probability per option
score2 to 10 ordered levels, lowest firstexpected level, which can land between two levels
noula yes/no questionprobability of yes, not a hard label

Questions are answered independently in llama.cpp's implementation. You can batch them: Kev-4B, lev, and OpenJev process the state only once. But don't expect the model to reason across questions.

The /v1/systemone API ​

Any client that already speaks System One works with llama.cpp after changing the base URL. Start a server, then send a state and your questions:

bash
llama serve -hf ggml-org/Kev-4B-GGUF:Q8_0
bash
curl http://localhost:8080/v1/systemone \
  -H "Content-Type: application/json" \
  -d '{
    "state": "Customer message: I was charged twice for my order last week and nobody has replied.",
    "questions": {
      "route": {
        "type": "choice",
        "instructions": "Which team should handle this?",
        "criteria": {
          "billing": "payments, charges, refunds, invoices",
          "shipping": "delivery, tracking, lost or late parcels",
          "technical": "bugs, errors, login problems"
        }
      },
      "angry": {"type": "noul", "instructions": "Is the customer angry?"},
      "urgency": {
        "type": "score",
        "instructions": "How urgent is this?",
        "criteria": ["can wait", "this week", "today", "right now"]
      }
    }
  }'

The response for that ticket: route billing at 0.90, angry at 0.82, urgency 2.28, which sits between "today" and "right now". Input tokens: 130. Output tokens: 0. That usage line is the whole pitch. In router mode, llama.cpp loads models on demand and you pick one per request with the model field.

Quick Take: Decision models don't replace chat models. They replace the JSON-mode, parser, retry-loop stack you were using for classification, routing, and gating decisions.

The models llama.cpp ships with ​

The collection started at five models. The PR thread's author immediately linked another one, with the note "looks like 5 is not enough". Add Cloudflare's Clef and you have six decision models in the System One family today.

ModelSizeBaseLanguagesImagesLicenseMedian per question*
Julia-1144MmmBERT-small50+noApache 2.03 ms
Laya421MModernBERT-largeEnglishnoApache 2.05 ms
Kev-4B4BQwen3.5-4B-BaseEnglishnoApache 2.012 ms
lev4BQwen3.5-4BEnglishnoApache 2.036 ms
OpenJev27BQwen3.8-27Ben, de, fr, hi, zh, jayesCC BY-NC 4.043 ms

*Median time to answer one question on one NVIDIA RTX PRO 6000.

Read the latency column as relative, not absolute. Julia-1 at 144M runs on a CPU all day. Kev-4B is small enough for a consumer GPU. OpenJev at 27B wants a workstation card, and the CC BY-NC license limits it to research. If you need to ship commercially, Kev-4B is the pick.

Clef: 27B, multimodal, Apache-2.0 ​

Clef is the biggest decision model with open weights and a permissive license. 27B parameters, Apache-2.0, built on Qwen3.8-27B including its vision encoder, so the state can be text, JSON, an image, or a video. The reference Python implementation defaults to a 16,384-token encoding limit, enough for a full contract or a long support thread in one pass.

The joint schema head is the architectural detail worth stealing. Instead of running an independent classifier per question, the head shares the backbone's evidence across all questions. Triage a message for route, urgency, and sentiment in one pass rather than three separate prompts.

The same weights also run as a regular multimodal chat model through vLLM or SGLang. The decision behavior comes from the System One request format plus the joint head, and the model card includes the Python code to load both. That dual personality is useful: build the decision pipeline, and you still have a normal VLM for the parts that need actual generation.

Where scoring beats generating (and where it doesn't) ​

Cloudflare published Decision Index results covering 40+ benchmarks. The pattern is not "bigger is better."

BenchmarkClefClef-flashJevKev 9BLaya
BFCL (case exact accuracy)98.598.895.894.538.1
BANKING77 (macro-F1)94.290.979.784.814.3
RouterBench (selected quality)79.779.979.980.057.1
MMLU (accuracy)90.391.891.775.330.7
GPQA Diamond (accuracy)48.051.078.338.827.6
CRUXEval (accuracy)86.786.173.051.240.2
Median latency (ms)209.338.8524.151.45.8

Clef cleans up on tool-use and structured tasks. 98.5 on BFCL means near-perfect on real API-call sequences. 94.2 macro-F1 on BANKING77 is production-grade intent classification. 86.7 on CRUXEval means it can reason about what code will output. It beats Jev on most of the index.

The reasoning-heavy rows tell the opposite story. GPQA Diamond goes to Jev at 78.3 against Clef's 48.0, and Jev also wins MMLU-Pro (82.7 vs 65.9) and BBH (92.9 vs 73.7). Graduate-level science questions still favor the generative approach. If your decision needs genuine multi-step reasoning, option scoring is a handicap, not a feature.

Latency changes which model you'd actually deploy. At 5.8 ms, Laya belongs inside a request path, on a CPU, no GPU needed. At 209 ms, Clef is an async review tool. At 524 ms, Jev is something you call when the answer is hard and you can wait.

Also notice Clef-flash, the small variant, beating the full 27B on several rows: BFCL 98.8 vs 98.5, GPQA 51.0 vs 48.0, Home appliance simulator 97.7 vs 83.0. On Typesafe's end-to-end workflow evals the big models cluster close together (invoice processing primary action: Clef 86.2, Jev 83.1; customer service exact: Clef 76.3, Jev 76.0). That's the healthy sign: the bottleneck is your options text, not the model.

What people are building with these ​

Three use cases keep showing up. Routing is the obvious one: support tickets, intents, document types. Moderation and risk scoring are second, since the noul and score types map directly to "does this need review" and "how bad is this on a scale". Third is agent observability, which is the one I find most interesting. The LocalLLaMA typed-decisions dataset is built around exactly this: given an agent's trace summary, decide whether it should continue, stop, be observed, or go to human review, and also score outcome, risk, and urgency. Each row carries agreement stats like argmax_agree and total_variation between scoring models. People are already using these to gate autonomous agents before they do something expensive or irreversible.

What the community is saying skews practical. When I first ran Julia-1 on a mock ticket, "I was charged twice" routed to shipping. Adding one-line descriptions to the options flipped it to billing at 0.99, so the label text is now the prompt. A vague "hi, quick question about my account" scored 0.25 with Julia-1 and 0.80 with Kev-4B, which broke the confidence cutoff I'd tuned for the smaller model. The advice that keeps surfacing in the llama.cpp announcement and the PR thread is consistent: try several sizes, write option descriptions, and tune your cutoff per model on your own examples.

Common Pitfalls ​

Five failure modes are worth knowing before you wire one of these into production.

Bare option labels. Julia-1 routed a double-charge complaint to shipping until each option had a description. At 144M parameters, labels like "billing" and "shipping" are not self-explanatory. Write criteria the way you'd write a grading rubric.

One confidence cutoff for every model. A vague ticket scored 0.25 with Julia-1 and 0.80 with Kev-4B. A cutoff that catches low-confidence misses on one model will flood your human queue on the other. Tune per model, on your own data.

Reading noul as a hard label. The noul answer is the probability of true. At 0.55, it's a maybe, not a yes. At 0.8, it's actionable. Treat it as a continuous signal and set your own threshold.

Rounding the score. An urgency score of 2.28 means between "today" and "right now". If you round to the nearest level, you discard the only information the model gave you. Use the fractional value.

License spillover. OpenJev is CC BY-NC 4.0. Free to experiment with, not free to ship commercially. If your project involves a payment processor, Kev-4B and Clef are the Apache-2.0 choices.

What's next ​

New open decision models are landing weekly, and llama.cpp says it will keep adding the best ones. The Decision Index leaderboard is becoming the reference table for comparing them. The trend I'd watch is efficiency: Clef-flash beats the 27B on several benchmarks while running five times faster. The frontier is moving down in size, not up. If you pick a model today, expect to re-pick in three months.

One Thing to Remember: the options are the interface now. You are not prompting a text model, you are defining the answer space. One-line descriptions instead of bare labels moved a 144M model from a wrong answer to a 0.99-confidence right answer. That's the difference between a decision system that works and one that quietly routes billing complaints to shipping.

The Bottom Line ​

If you're building a router, moderation filter, or agent-safety gate, adopt a decision model now. Kev-4B gives you 4B-class accuracy at 12 ms per decision under Apache-2.0, and you keep your chat model for the parts that need generation.

If your input is documents, screenshots, or video, test Clef first. Its 16K-token context and vision cover a full page in one pass, and at 209 ms median latency it's priced for review queues. For the lowest-latency CPU path, Laya at 5.8 ms wins by default.

One thing to watch: the model zoo is turning over weekly, and small models are already beating their bigger siblings on real benchmarks. Expect "the right model for routing" to be a moving target for the rest of the year. Keep the System One API as your stable layer and let the model underneath rotate.