Mingbird: Small Models Fail Because of the Harness, Not the Model
Put a 4B model inside a harness built for cloud frontier models and it falls apart. Tool definitions overflow the context, self-correction spirals, demos loop, tasks get quietly abandoned. Mingbird (arXiv 2610.02001) argues most of that is the harness's fault, and builds one for small models to prove it.
It's a local-first harness for Windows and Ollama with ten mechanisms, each patching a specific small-model failure. Three stand out: a byte-level prefill budget so tool definitions never eat the context, a finish gate that re-reads the task before accepting "done", and loop detection by action signature. On the authors' LRAB benchmark (4 harnesses x 4 open models from 2B to 35B x 18 real tasks, all 288 cells published) Mingbird scores 0.886 against 0.631 for goose, 0.479 for opencode and 0.405 for agent-mini. On tau2-bench it scores 0.856 against 0.791 and 0.737.
The frontier-model probe is the most interesting line. On the same 18 tasks a frontier model ranges from 0.997 to 0.478 depending on harness. Well-formed scaffolds stay within 0.072 of each other. A bad harness can halve a top model.
The authors are honest about limits: self-built benchmark, one machine, single-trial scoring, and same-night reruns move scores by up to 0.069, as large as most single-mechanism effects. Still, the finish gate alone gives a paired +0.10 across three replications. "Re-read the task before you say you're done" is a one-line idea anyone can steal for their own agent today.
Link: arxiv.org/abs/2610.02001
← Back to all articles
It's a local-first harness for Windows and Ollama with ten mechanisms, each patching a specific small-model failure. Three stand out: a byte-level prefill budget so tool definitions never eat the context, a finish gate that re-reads the task before accepting "done", and loop detection by action signature. On the authors' LRAB benchmark (4 harnesses x 4 open models from 2B to 35B x 18 real tasks, all 288 cells published) Mingbird scores 0.886 against 0.631 for goose, 0.479 for opencode and 0.405 for agent-mini. On tau2-bench it scores 0.856 against 0.791 and 0.737.
The frontier-model probe is the most interesting line. On the same 18 tasks a frontier model ranges from 0.997 to 0.478 depending on harness. Well-formed scaffolds stay within 0.072 of each other. A bad harness can halve a top model.
The authors are honest about limits: self-built benchmark, one machine, single-trial scoring, and same-night reruns move scores by up to 0.069, as large as most single-mechanism effects. Still, the finish gate alone gives a paired +0.10 across three replications. "Re-read the task before you say you're done" is a one-line idea anyone can steal for their own agent today.
Link: arxiv.org/abs/2610.02001
Comments