DevOps pipelines pay for LLM generation even when they only need a yes, no or label. Decision models offer cheaper classification while keeping thresholds and gates in code. The hosts explore Jev and eight candidate jobs, from alert triage and duplicate detection to PII checks, model routing and GraphRAG. They explain where classifiers fit around agent tool calls, how to route uncertain probabilities for review, and why focused input and deterministic checks still matter.
Summary
A lot of what you send an LLM today doesn’t need a sentence back. Is this alert noise? Is this ticket a duplicate? Does this log line contain PII? Does this tool call match what the user asked for? For questions like these you’ve been paying a frontier model to reason and generate, even when a schema keeps the output tidy, and then waiting while it produces the answer one token at a time. Decision models, also called System One models, are built for exactly this shape of question. You give them a state and a set of typed questions. They return probabilities and choices in one pass, for a small fraction of a cent. Our position is to treat one as “a smart if statement”. Use it to take a reading from an input. Keep every threshold and every gate in your own code. Don’t let it decide anything that needs more context than its small window holds, and don’t treat it as a security boundary.
What a decision model returns that an LLM doesn’t
Structured answers aren’t the new part. Speed, price and probabilities are.
How is a decision model different from an LLM with structured output?
It doesn’t generate. It answers every question you send in a single pass and returns a probability, not a stream of tokens shaped to look like JSON.
Schema-constrained output from LLMs is a solved problem. Anthropic’s structured outputs use constrained decoding to keep a response inside your schema, and in episode #20 we recommended them over the old prefill trick. So you don’t have to parse prose to get a yes or no out of Claude. What you still pay for is an autoregressive model writing its answer token by token, and with a thinking model, reasoning at length first. A decision model skips the generation. TypeSafe’s announcement describes Jev as producing “all outputs in a single query”.
Take a support ticket. You send the ticket text as the state, plus the questions. Is it urgent? Is the customer satisfied? Does a person need to get involved? Is it billing or technical support? Each one comes back with a number attached. A yes/no question, which TypeSafe calls a noul, returns the probability that the answer is yes. Choice questions pick from options you supply, and score questions place the input on a scale you describe. Questions about the same state are answered together, which makes several questions in one request much cheaper than several separate LLM calls.
Price is the other half. TypeSafe’s model documentation lists Jev input at $42 per billion tokens, which is $0.042 per million, and output is free. We sent about 50 questions while trying it, and the bill came to less than a cent. Five dollars goes a long way.
On speed, TypeSafe quotes 70 to 500 ms end to end. Cloudflare’s own benchmark measured a 524 ms median for Jev against 38.8 ms for its smaller Clef-flash model. That’s a vendor benchmark. Both numbers are fast for a model call, but half a second is still a lot to add to a synchronous request path. Whether you can afford to run a check on every event, or only on a sample, depends on your volume, your concurrency and the rate limits you hit. TypeSafe currently lists 40 requests per second, subject to change. Measure it on your own workload before you put it inline.
Where it fits in an agent loop: hooks before and after tool calls
If the model is fast and cheap but can’t reason deeply, the best place for it is somewhere a quick reading saves an expensive step. In an agent harness, that’s the hooks around tool calls.
Where does a decision model earn its keep in an agent harness?
In pre-tool and post-tool hooks. It catches honest mistakes before you spend time and tokens on them, and it filters which results are worth learning from.
Before the tool runs, the agent says “I want to call this, with these arguments.” Start with the checks that have a definite answer. Does the call match the tool’s schema? Does the resource exist? Is the flag supported? Do those in code. Validate against the schema and look the resource up in your source of truth, the inventory or the API itself. A classifier can’t tell from the arguments alone whether prod-db-2 exists or whether a CLI accepts --force-all, and it only knows what you put in its state. Leave the semantic judgment to the decision model: given the user’s request and the proposed arguments, does this call look like what was asked for? If any check fails, the hook returns “wrong arguments” right away. You don’t run the tool, wait for it to fail, and feed the error back through the LLM. We wrote about hooks as the place to enforce rules in episode #5, and this is a cheap new check to put in them.
A shell command works well as input. Commands are short, so even a 1,000-token context is enough. One of our setups runs a Python script that asks a general LLM “is this command safe?” and also asks a local decision model, as a separate guardrail, “what’s the chance this command kills my laptop?” Its answer comes back almost instantly and flags the obviously bad cases.
After the tool runs you get a second chance. Every tool result is a bit of experience. Was the call successful? Was the output useful? A post-tool hook can ask a decision model whether this result is worth learning from. Only when the answer is yes do you pay a general-purpose LLM to write a learning snippet for the agent’s memory. The expensive step runs only when it’s worth it, which matches the memory hygiene we argued for in episode #7.
Treat it as a first filter for mistakes. Your real permissions and policy checks still have to stand on their own.
Keep the decision in your code
A hook that returns “wrong arguments” is a small decision. The temptation is to give the same model bigger ones.
Can a decision model gate a deployment?
Not on its own. Let it give you a reading, and keep the thresholds and the actual go or no-go in code you own and can review.
It’s tempting to point a score at a blue-green switch, a deployment gate or a router and let it decide. Don’t. Use the model as a classifier, a way to get a reading from information, and keep the guardrail logic in code. “Contains PII: 0.87” is an input to your logic. Your code decides what 0.87 means for this log stream, and that threshold can be reviewed, tested and changed in a pull request. Ask for the probability rather than a bare yes or no, so your code gets to choose the cut-off.
Read the number for what it is. A noul returns one probability that the answer is yes. Near 1 is a confident yes. Near 0 is a confident no. Around 0.5 is where the model can’t tell, as TypeSafe’s documentation explains, and there’s no separate confidence value. So “a low score goes to a human” gets it backwards. A 0.02 on “does this alert need a person?” is a confident no and can take the automated path. Route by distance from the middle instead. TypeSafe’s own example sends anything between 0.2 and 0.8 to review. A clear yes or a clear no goes to whichever automated path your code assigns, and the uncertain middle goes to an LLM that can reason about it, or to a human when acting on a wrong answer is expensive. Choice and score questions return different shapes, so check what each question type’s number means before you reuse the bands.
Then validate the bands. Run a set of representative inputs you’ve already labelled, such as last month’s alerts with the outcome known, and see where the wrong answers land. Move the threshold up where a false yes is expensive, like paging someone, and down where a missed yes is expensive, like a leaked secret.
Context size is the second reason to keep big decisions away from it. Jev allows 32K tokens for the state plus the longest question. Some local models we’ve run take around 1,000. A blue-green decision depends on metrics, recent deploys, error budgets and the dependency graph. You can’t fit that into a tiny window and expect a good call. If a question needs reasoning because it’s complicated, this is the wrong tool. That’s the “smart if statement”: you can pack a lot of conditions into it, but it isn’t there to decide for you.
DevOps jobs worth trying first
Keep the decision in code, keep the question small, and a long list of candidate jobs opens up.
Which DevOps problems are a good first experiment?
High-volume, low-complexity questions where you’d have liked to use an LLM but the cost or latency made it impractical.
Most of the launch demos classified email inboxes. These are the ones we’d try first:
- Alert and security-finding triage. Score each alert and send the noise one way and the likely real issues another. Early reports describe people filtering security-issue noise this way, which ties into the alert-fatigue problem from episode #19.
- Duplicate issues and PRs. Keep issue and PR descriptions, with code snippets, stored locally. When a new one arrives, compare it against all 2,000. We’d expect the decisions to finish before the same 2,000 items would come back through the GitHub API, but that’s an expectation, not a measurement. Time it on your own backlog, with your own concurrency and rate limits.
- Clustering, then naming. A decision model can tell you a group of issues is about the same thing, but it can’t tell you what that thing is. Send each cluster to a general-purpose LLM and ask it to name it. The small model handles the volume and the big model writes the label.
- PII detection in logs. Ask for a probability, not a bare yes or no, and let your code set the threshold.
- Model routing and token budgets. Is this task easy enough for a cheap model? Is the question worth spending tokens on at all?
- Deterministic versus generative. If a predefined script very likely solves the problem, run the script. If not, call the LLM. The same check works inside a coding agent: generate new code, or reuse an internal tool the organisation already has?
- Graph building. For GraphRAG ingestion, people call an LLM to propose links between pieces of information, which is slow and expensive. With a predefined set of relationship types, a decision model can answer “do these two connect?” and then “which of these ten relationships is it?”. Graph-RAG was the subject of episode #11.
- CAPEX versus OPEX. On the management side: is this work building something new you can sell, or maintaining what you already have?
These models also handle changing state well. You pass the current state, get a decision, the state changes, and you pass the updated state. That fits agent loops and incident timelines, where the picture keeps moving.
The state is the new prompt
Every job on that list depends on one input, and getting it right is most of the work.
What limits quality in practice?
What you put in the state. Unrelated material makes the answers worse, and the context window caps how much you can include.
These models call their input the state. Choosing what goes into it is the same engineering problem as writing a prompt, aimed at a different model. Include only what the question needs: one command, one log line, one issue and its nearest neighbours. You’ll hit the context limit as soon as you try anything bigger, and it’s worth learning that early. Larger windows are the obvious next improvement, and Cloudflare’s Clef already offers 64K.
Often the hard part is picking the first use case. You have a new toy, so what do you use it for? Look for places in your stack where the real question is a decision and you’re paying for generation, or where nobody asks at all because asking cost too much.
Placement is flexible. The smaller open-weight models can run on a CPU next to a local agent, but hardware needs vary a lot by model. AWS sized Strands Decider 2B for consumer hardware, while Clef ships in 9B and 27B sizes. Check the memory footprint of the model and quantization you actually pick before you plan around a laptop. In the cloud, you can self-host one as an ECS service or on SageMaker and call it from your applications. Self-hosting also changes the latency picture from the first section: no network hop to a vendor, but your own hardware sets the throughput.
Not a new idea, and copied in just over two weeks
All of this matters for a practical reason: you aren’t locked into one vendor to try it.
Is this really a new kind of model?
As a product category, yes. As an idea, no. Classification is old machine learning, and the reaction to Jev showed how easy the approach is to copy.
Ten-something years ago, classifying email as spam or not spam took a fairly complex piece of software. A tiny model can now do the same job. TypeSafe AI announced Jev in early access on 15 September 2026, after reportedly spending around two years on it. Just over two weeks later the field was crowded:
- OpenAI announced a Decisions API at DevDay at the end of September. Its DevDay recap describes it as running on GPT-6 Luna over user-defined questions with predefined answers, in limited preview.
- Cloudflare released Clef on 1 October, in 27B and 9B sizes on Workers AI. It accepts Jev’s request format, ships Apache 2.0 weights and has a 64K context window.
- AWS released Strands Decider 2B the same day. It’s a small Qwen-based model with its text head replaced by a scoring head, sized for consumer hardware.
- Individuals built their own versions on Qwen bases and published them.
If we’d built Jev, we’d feel a bit upset. You let the genie out, and everyone does the same thing a couple of weeks later. One builder made the point loudly. The author of Laya published a similar non-autoregressive approach in March 2025. It was trained on synthetic sales-call transcripts to coach a salesperson in real time, and almost nobody noticed. His complaint was that well-funded general-purpose launches get the attention while narrow, useful work gets none. It moves fast. It is what it is.
For teams buying, the copycats are good news. While waiting on the Jev waitlist, we got an open-source alternative running locally before access arrived. That leaves TypeSafe with a question to answer. If you can run the thing on a laptop and inference is nearly free, what do you pay for? Private hosting, or custom fine-tuning for your own decisions, are plausible answers. We’d wait and see before building a dependency on one vendor’s API. Several of the alternatives accept the same request format, which makes switching cheap.
Beyond the hype: the right model for the job
The speed of the copying says something about where models are heading.
Is this the start of something bigger?
Probably. It’s renewed interest in specialised models for the decisions inside agent workflows, after a stretch where building with AI mostly meant calling a large language model.
One commenter compared a decision model next to a frontier LLM to a bicycle next to a space station. Your bicycle is fast, cheap and simple, but it does one thing. That’s the point, though. For teams building agents, the recent wave has been LLMs, LLMs, LLMs. This is a different shape of model, and it says plainly that not every step in an agent needs an LLM. Specialised models may follow. A coding model could emit a whole function at once instead of one token at a time. Eventually, choosing a model would mean choosing the type of intelligence that fits the problem, not just choosing between Anthropic and OpenAI.
Apple seems well placed. Its Foundation Models framework already gives apps on-device inference on Apple silicon. A decision model small enough to run on a phone could bring mail classification to Mail and triage to Notes without installing anything else. Local LLMs make laptops run hot. A model that never generates text might not.
What this means for teams
Whether the hype dies down or something durable comes out of it, finding out is cheap. Four things to try this week:
- Find one decision you’re paying generation for. Search your agent hooks and pipelines for LLM calls whose output is a yes, a no or a label. Pick the highest-volume one.
- Run an open-weight decision model locally against it. Put only what the question needs in the state. Note where you hit the context limit.
- Label a sample and set the bands. Take last month’s alerts, issues or log lines with the outcome known, run them through, and choose thresholds by distance from 0.5. Put those thresholds in code and review them in a pull request.
- Wire it into a pre-tool hook behind your deterministic checks. Schema validation and resource lookups first, the semantic check after. Measure latency end to end before it goes anywhere near a synchronous path.
If you build something we didn’t list, tell us on LinkedIn.
Common questions, answered
How do I set thresholds for a decision model’s yes/no probability?
Route by distance from 0.5, not by how high the number is. A yes/no probability near 1 is a confident yes, near 0 a confident no, and around 0.5 is uncertain. TypeSafe’s documentation for Jev uses 0.2 and 0.8 as an example, sending everything in between to review. Keep those thresholds in your own code, check them against labelled examples from your own data, and move them according to which mistake costs more.
Can I use a small classifier model as a guardrail for AI agent tool calls?
Yes, for catching honest mistakes, but not as a security control. Put it in a pre-tool hook after deterministic checks: validate the arguments against the tool’s schema and look up resources in your source of truth first. Then ask the classifier semantic questions, such as whether the call matches what the user asked for or whether a short shell command looks destructive. Someone with malicious intent can get input past a classifier, so permissions and policy checks still have to stand on their own.
How do I find duplicate GitHub issues cheaply with a local model?
Store issue and PR descriptions locally and ask a decision model whether each new issue matches an existing one. Comparing a new issue against a couple of thousand stored ones is cheap with a small open-weight model on a CPU, but benchmark it on your own backlog before relying on it. To label the clusters it finds, send each group to a general-purpose LLM and ask it to name the common topic.
Related episodes
- Episode #5: Stop Your Agent Before It Breaks Prod: the hooks and guardrails that a pre-tool-use decision check plugs into.
- Episode #7: When Agent Memory Helps and When It Hurts: why a post-tool filter on what gets written to memory matters.
- Episode #11: Base of Record for Intelligent Systems: the graph-RAG background for using decision models to propose links between pieces of information.
- Episode #19: AI Doom Can Wait, Your Security Alerts Can’t: the alert-fatigue problem that cheap triage scores are aimed at.
- Episode #20: Become an Agentic Engineer: structured outputs for LLMs, the baseline a decision model has to beat.
Resources
- Introducing System One models and Jev (TypeSafe AI): the launch post. Settles the 15 September early-access date, the single-pass design and TypeSafe’s own 70 to 500 ms latency claim.
- Jev models and Noul (TypeSafe docs): the primary source for pricing, rate limits and the 32K state limit, and for how to read a yes/no probability and set review bands.
- Structured outputs (Claude API docs): the LLM baseline. Schema compliance through constrained decoding, which is why the case for a decision model rests on speed, price and probabilities.
- Introducing Clef (Cloudflare changelog): open weights under Apache 2.0, Jev-compatible requests and a 64K window, plus the latency comparison above. Treat the benchmark as a vendor’s. The 9B and 27B sizes are a reminder that “runs locally” depends on which model you pick.
- Introducing Strands Decider 2B (Strands Agents): AWS’s small open model for consumer hardware. Its intended uses (model routing, tool selection, guardrails, memory management) line up with the hook patterns above.
- DevDay 2026 recap (OpenAI): OpenAI’s own description of the Decisions API on GPT-6 Luna and its limited-preview status. Check access before you design around it.
- Laya and SalesRLAgent (arXiv:2503.23303): the open-source model and the 2025 paper behind the “I did this a year ago” story. Worth reading if you want a narrow, domain-trained example.
- Apple Foundation Models framework: the on-device inference API behind the hope for a classifier built into Mail.