October 2, 2026ResearchBenchmarkMonitoring

Jev-Style Decision Models Quietly Squash Your Rating Scale, Study Finds

On the same day Cloudflare and Amazon shipped decision models, a paper showed the failure mode that matters most for them. "More Choices, Fewer Decisions" (arXiv 2609.38827, 49 upvotes on Hugging Face) tested JEV 1.13 and three open KEV models. It looked at whether they actually use the decision scale you give them. Often they do not.

The paper opens with one striking case. On ANLI, a three-way task (entailment, neutral, contradiction), Jev is 74.95% accurate. But it puts 38.8% of all predictions, and 51.3% of its errors, into Neutral, even though the gold labels are nearly balanced and candidate positions were balanced too. The model is not bad. It just hedges toward the middle.

Then it scales up. Across 36 ordinal datasets, the final decisions use only 67–76% of the effective range of the gold labels. On four nominal tasks with unordered categories the figure is 87–102%, which is fine. So the compression is specifically about ordered scales: ratings, severities, likelihoods. Make the scale finer and it gets worse. Going from 2 options to 14, utilization falls for every model and lands at 26–75% at K=14. The odd part is that the candidate probabilities stay broad for most models. The model "knows" the extremes are plausible, and the decision step still collapses toward the center. Shuffling candidate order helps a little but does not remove it.

The good news is that it can be trained out. Targeted BA-LoRA post-training lifts utilization from roughly 47% to 86% on eight supervised scales. So this is a learned habit, not something built into the architecture.

Why it matters this week: the killer use for these models is cheap per-action review of agents, and review is almost always ordinal. Low, medium or high risk. Allow, flag or block. A judge that drifts to "medium" and "flag" lets the scale's ends quietly die. If you use a decision model as a grader or monitor, plot your output distribution against your labels before trusting the accuracy number.

Link: arxiv.org/abs/2609.38827
← Previous
MILO Evolves Its Own Harness and Tops Terminal-Bench With 26% Fewer Tokens
Next →
Impeccable: 61 Rules That Catch AI-Generated Frontend Design
← Back to all articles

Comments

Loading...
>_