The Decision-Model Savings Were 23.9%. After the Authors Audited Themselves, 4.3%.
Decision models are this fall's favorite harness trick. A small System-1 model answers the many tiny typed questions an agent harness asks per task (which model to call, which tool to use, is this retrieved text relevant, does this input carry an injection) in one forward pass, with class probabilities, for a fraction of an LLM call. A new arXiv paper runs the first careful head-to-head, and then does something rarer. It audits its own pipeline and publishes what it got wrong.
The comparison is an open-weight model, Laya, against the hosted Jev, on 11 agent decision points built from 18 public sources: 7,283 base cases plus 6,640 robustness variants, with byte-identical inputs and reproducibility checks across hardware and across days. Jev is significantly more accurate on 9 of 11, by 10.8 to 46.0 points. Laya changes 30% of its answers when the option order is reversed, and with 50 similar tools to pick from it drops to 31% while Jev holds 98%.
Two results cut against the hype for both. Neither model beats chance at zero-shot model routing, which is one of the most advertised use cases. And they tie on RAG relevance gating.
Then comes the self-audit. Three analysis errors and one design confound had inflated the deployment claims. A pre-screen cost had been left out, so a reported 23.9% saving was really 4.3%. Gate accuracy had been reported as if it were end-to-end quality, 98% against an actual 58%. Thresholds had been tuned in-sample, so a 5% miss target became up to 17% on held-out data. And a "channel effect" on injection false positives disappeared once the test content was native to the channel.
That last section is the reason to read it. Every one of those mistakes is easy to make in a production rollout: counting the cheap call but not the pre-screen in front of it, quoting the gate's accuracy as the system's accuracy, tuning the threshold on the data you report. If a research team with paired tests fell into all four, a vendor case study probably did too. The practical rule that follows is to ask for end-to-end numbers on held-out traffic, with every stage's cost included.
Link: arxiv.org/abs/2610.02267
← Back to all articles
The comparison is an open-weight model, Laya, against the hosted Jev, on 11 agent decision points built from 18 public sources: 7,283 base cases plus 6,640 robustness variants, with byte-identical inputs and reproducibility checks across hardware and across days. Jev is significantly more accurate on 9 of 11, by 10.8 to 46.0 points. Laya changes 30% of its answers when the option order is reversed, and with 50 similar tools to pick from it drops to 31% while Jev holds 98%.
Two results cut against the hype for both. Neither model beats chance at zero-shot model routing, which is one of the most advertised use cases. And they tie on RAG relevance gating.
Then comes the self-audit. Three analysis errors and one design confound had inflated the deployment claims. A pre-screen cost had been left out, so a reported 23.9% saving was really 4.3%. Gate accuracy had been reported as if it were end-to-end quality, 98% against an actual 58%. Thresholds had been tuned in-sample, so a 5% miss target became up to 17% on held-out data. And a "channel effect" on injection false positives disappeared once the test content was native to the channel.
That last section is the reason to read it. Every one of those mistakes is easy to make in a production rollout: counting the cheap call but not the pre-screen in front of it, quoting the gate's accuracy as the system's accuracy, tuning the threshold on the data you report. If a research team with paired tests fell into all four, a vendor case study probably did too. The practical rule that follows is to ask for end-to-end numbers on held-out traffic, with every stage's cost included.
Link: arxiv.org/abs/2610.02267
Comments