September 28, 2026InfrastructureCodingRL

Ember-1: Fireworks Taught Kimi K3 to Shut Up and Think

Thinking models think too much. Fireworks says Kimi K3 spends sometimes more than 90% of its output tokens on internal reasoning, and in an agent loop that bill compounds, because every turn replays the earlier reasoning back into context. Ember-1 is Fireworks' answer: a model built on K3 and trained to reason in fewer tokens, with the same answers.

The numbers are the kind that make procurement people sit up. On two customers' production coding traffic, about 35% fewer tokens per task at comparable quality. In one of those customer comparisons, reasoning tokens dropped 71% and total tokens 39%, while steps per task fell from 23.8 to 21.4 and the score went slightly up. On benchmarks priced at K3's public API rates, Ember-1 matches or beats K3 at max effort: 82.0% vs 80.9% on Terminal Bench 2.1, 92.2% vs 93.2% on SWE-bench Verified, and 75.2% vs 66.4% on DeepSWE 1.1, all at a fraction of the cost. One customer has already moved it into production and plans to replace the base model entirely.

The most useful sentence in the post is about what did not work. Turning K3's reasoning effort down gave up too much quality. The low setting is strictly dominated by Ember-1 on their charts. So the lever everyone reaches for first, the effort knob, is the wrong one. Shorter reasoning has to be trained in, with feedback from real tasks, including teaching the model to give up sooner on attempts that are going nowhere. Fireworks says it took 50-plus training runs and 200-plus evals, all on its own serverless training platform.

Read this next to last week's price cuts from OpenAI and Anthropic and the direction is obvious. The market has stopped asking which model is smartest and started asking what a completed task costs. Ember-1 attacks that number from the side nobody prices: not cheaper tokens, fewer of them. It also says something about where inference companies are heading. Fireworks is no longer just serving other people's weights. It is shipping its own post-trained variants, and a series is promised. The blog post is dated September 23; it reached the Hacker News front page on Sunday.

Link: fireworks.ai/blog/ember-1
← Previous
Tens of Thousands of Incidents, One Training Pause, and a Fight Over the Word Rogue
Next β†’
OpenRig: A Harness Wraps a Model. A Rig Wraps Your Harnesses
← Back to all articles

Comments

Loading...
>_