Fireworks AI Releases Ember-1: Kimi K3 with 40% Fewer Tokens
Fireworks AI launches Ember-1, a post-trained Kimi K3 model available via serverless API, cutting token usage by 35 to 50 percent.

Stock photo for illustration only, not from the actual event
- Fireworks AI has introduced Ember-1, a post-trained version of Kimi K3 available as a research preview.
- The model reduces internal reasoning token usage by 35% to 50% without sacrificing accuracy.
- Live A/B testing with production coding workloads lowered total output tokens from 49.3K to 29.9K per task.
- It shares the same pricing tier as Kimi K3, generating cost savings purely through higher token efficiency.
Fireworks AI has officially released Ember-1, a newly post-trained artificial intelligence model built upon the Kimi K3 foundation. Designed to address high operational costs in multi-turn agentic workloads, the model tackles the issue where reasoning models frequently spend over 90% of generated tokens purely on internal deliberation, compounding expenses quadratically across turns.
Previously, attempts to lower Kimi K3's reasoning effort compromised overall task quality. To resolve this without sacrificing capability, the engineering team trained the model to reason more efficiently, preserving constructive self-reflection while eliminating redundant loops and unproductive reasoning cycles.

Stock photo for illustration only, not from the actual event
The training regimen spanned diverse domains including mathematics, coding, instruction following, conversation, search, tool use, and software engineering. Spanning over 50 training experiments and 200 evaluations executed entirely on Fireworks Serverless Training infrastructure, the model was trained using proprietary data without incorporating any customer information.
The economic burden of internal reasoning tokens represents a major bottleneck for enterprise AI deployment, particularly in multi-step agent frameworks. By optimizing the reasoning trajectory rather than merely dampening effort levels, Fireworks AI demonstrates a scalable path toward making sophisticated agentic workflows economically viable in production environments.
In live A/B tests conducted with two enterprise customers on production coding tasks, output tokens dropped from 49.3K to 29.9K per task. Specifically, reasoning tokens fell by 71.3% while overall token consumption decreased by 39%. Crucially, task performance remained steady, scoring 0.753 for Ember-1 compared to 0.751 for the baseline Kimi K3 model.
Ember-1 is priced identically to Kimi K3 on the Fireworks platform, charging $3.00 for input, $0.30 for cached input, and $15.00 for output per 1 million tokens. Financial savings are achieved entirely through reduced volume generation, and at least one customer has already integrated Ember-1 into their active production pipeline.
Source: MarkTechPost
Found something wrong in this article? Report an issue with this article
Comments
Leave a Comment