Clef, Strands Decider, Decisions API: Three Giants Clone Jev in One Week
A month ago "decision model" was one startup's pitch. This week it became a product category. OpenAI announced a Decisions API at Dev Day on Tuesday. On Thursday, Cloudflare released Clef and Clef-flash, and AWS released Strands Decider 2B. All three do the same thing TypeSafe's Jev does: you hand the model a fixed set of options, and it hands back a typed choice with calibrated probabilities. No free text and no tool-call improvisation.
Cloudflare's entry is the most serious of the three. Clef runs on a frozen Qwen3.8-27B and Clef-flash on a frozen Qwen3.5-9B, each with a routing head and rank-256 adapters on top. The model does one prefill pass and then scores every valid schema choice in parallel. Nothing is generated token by token, which is why it is fast. It reads images, which Jev does not, and has a 64k context against Jev's 32k. Cloudflare says it beats Jev in 3 of 4 areas on TypeSafe's own eval suite. The weights are Apache 2.0 on Hugging Face, the API is Jev-compatible, and hosting is on Workers AI. Cloudflare's own threat-intel team used it to classify a domain in 2.2 seconds, against 4.7 seconds for gpt-oss-120b, and got more categories back. An RL fine-tuning service starts as hands-on work with Cloudflare engineers and becomes self-serve later.
Amazon's version started as a side project. Distinguished engineer Marc Brooker built it on Qwen3.5-2B after seeing Jev. It briefly topped Jevbench for its size, so Strands Labs cleaned it up and open-sourced it. Brooker's pitch is the cleanest summary of the whole category: a perfect decider for a workflow step, the "what do I do next, given where I am" call.
The number that explains the rush comes from a hackathon demo, not a lab. It used Jev to check every agent action against the original task. That cost $2.94, versus $372 with a frontier LLM. At that price you can afford a reviewer on every single action, and that is exactly the agent-monitoring problem OpenAI has been paying "significant compute cost" to solve.
TypeSafe's CEO Diogo Almeida has the best line on all this: "If you want it really fast and cheap, use dice." The open question is calibration, not speed. Training one of these costs hundreds to thousands of dollars, so nobody has a moat yet. The first paper to stress-test calibration on ordinal scales has already found a real flaw, and it is covered separately today. Expect a dozen more of these models by November. Pick yours by its calibration curve, not its latency chart.
Links: blog.cloudflare.com/clef-decision-models, techcrunch.com (Strands Decider 2B, Oct 1)
← Back to all articles
Cloudflare's entry is the most serious of the three. Clef runs on a frozen Qwen3.8-27B and Clef-flash on a frozen Qwen3.5-9B, each with a routing head and rank-256 adapters on top. The model does one prefill pass and then scores every valid schema choice in parallel. Nothing is generated token by token, which is why it is fast. It reads images, which Jev does not, and has a 64k context against Jev's 32k. Cloudflare says it beats Jev in 3 of 4 areas on TypeSafe's own eval suite. The weights are Apache 2.0 on Hugging Face, the API is Jev-compatible, and hosting is on Workers AI. Cloudflare's own threat-intel team used it to classify a domain in 2.2 seconds, against 4.7 seconds for gpt-oss-120b, and got more categories back. An RL fine-tuning service starts as hands-on work with Cloudflare engineers and becomes self-serve later.
Amazon's version started as a side project. Distinguished engineer Marc Brooker built it on Qwen3.5-2B after seeing Jev. It briefly topped Jevbench for its size, so Strands Labs cleaned it up and open-sourced it. Brooker's pitch is the cleanest summary of the whole category: a perfect decider for a workflow step, the "what do I do next, given where I am" call.
The number that explains the rush comes from a hackathon demo, not a lab. It used Jev to check every agent action against the original task. That cost $2.94, versus $372 with a frontier LLM. At that price you can afford a reviewer on every single action, and that is exactly the agent-monitoring problem OpenAI has been paying "significant compute cost" to solve.
TypeSafe's CEO Diogo Almeida has the best line on all this: "If you want it really fast and cheap, use dice." The open question is calibration, not speed. Training one of these costs hundreds to thousands of dollars, so nobody has a moat yet. The first paper to stress-test calibration on ordinal scales has already found a real flaw, and it is covered separately today. Expect a dozen more of these models by November. Pick yours by its calibration curve, not its latency chart.
Links: blog.cloudflare.com/clef-decision-models, techcrunch.com (Strands Decider 2B, Oct 1)
Comments