SearchJev: A Decision Model for Search Agents, 5x Faster With Better Answers
Search agents make the same short decisions over and over. Is this result relevant. Is the evidence enough. Should I search again or answer. Using a generative model for each one adds latency and gives you confidence numbers nobody should trust. A paper posted Monday, SearchJev, pulls those decisions out into a separate fast model and leaves the slow model to plan, write queries and compose answers. The name is deliberate: it is the Jev decision-model pattern applied to one agent loop.
Given a search state and a decision schema, SearchJev scores the legal options directly, with no autoregressive generation. The training trick is Soft-Label Learning for Calibrated Decisions, which learns decision probabilities from uncertain supervision and calibrates the confidence. The uncertain judgments get escalated to the System-2 model. The authors also release SearchDecision-Bench, a benchmark unifying six kinds of search decision for training and evaluation.
The numbers are the kind the decision-model thread has been waiting for. On SearchDecision-Bench, SearchJev beats same-size Qwen3.5 autoregressive models on decision quality, makes decisions 5.2 to 5.3 times faster, and cuts expected calibration error by 41 to 74%. On BrowseComp-Plus the dual-system agent gets a 3.7 to 4.7 times speedup in active search time while answer accuracy rises from 45% to as much as 54%. Faster and more accurate at the same time is the result that the ordinal-bias and zero-shot-routing critiques from the past week said was hard to get. The difference is that this one is trained on the exact decision schema it serves.
The design pattern generalizes past search. Any agent loop has a handful of yes/no gates that fire hundreds of times per task and one or two generation steps that matter. Pricing the gates at classifier cost and the generation at LLM cost is the cost-per-task argument in its cleanest form.
Link: arxiv.org/abs/2610.05107
← Back to all articles
Given a search state and a decision schema, SearchJev scores the legal options directly, with no autoregressive generation. The training trick is Soft-Label Learning for Calibrated Decisions, which learns decision probabilities from uncertain supervision and calibrates the confidence. The uncertain judgments get escalated to the System-2 model. The authors also release SearchDecision-Bench, a benchmark unifying six kinds of search decision for training and evaluation.
The numbers are the kind the decision-model thread has been waiting for. On SearchDecision-Bench, SearchJev beats same-size Qwen3.5 autoregressive models on decision quality, makes decisions 5.2 to 5.3 times faster, and cuts expected calibration error by 41 to 74%. On BrowseComp-Plus the dual-system agent gets a 3.7 to 4.7 times speedup in active search time while answer accuracy rises from 45% to as much as 54%. Faster and more accurate at the same time is the result that the ordinal-bias and zero-shot-routing critiques from the past week said was hard to get. The difference is that this one is trained on the exact decision schema it serves.
The design pattern generalizes past search. Any agent loop has a handful of yes/no gates that fire hundreds of times per task and one or two generation steps that matter. Pricing the gates at classifier cost and the generation at LLM cost is the cost-per-task argument in its cleanest form.
Link: arxiv.org/abs/2610.05107
Comments