Cohere Releases North Small Translate 218B MoE Model
Cohere introduces North Small Translate, a 218B MoE translation model scoring 83.6 on WMT26 across 50 languages with an Agentic variant.

Stock photo for illustration only, not from the actual event
- North Small Translate is the first translation model in Cohere's North product family.
- It uses a decoder-only sparse MoE architecture with 218B total and 25B active parameters.
- The model achieves an average score of 83.6 on WMT26 benchmarks across 50 languages.
Cohere has officially released North Small Translate, a decoder-only sparse Mixture of Experts (MoE) translation model featuring a total of 218 billion parameters. Approximately 11.5% of its weights are active per token, meaning per-token compute tracks the 25 billion active parameters while its memory footprint requires holding all 218B weights.
This model marks the first translation release within Cohere's North family, succeeding earlier multilingual lineages such as Tiny Aya and Command A Translate. Cohere partnered with RWS and its Language Weaver scientists and language experts to refine the model's real-world translation quality, emphasizing that global communication is foundational to organizational sovereignty.

Stock photo for illustration only, not from the actual event
In terms of performance, the model reports a WMT26 all-languages score of 83.6 across 50 supported languages. It also features an Agentic variant capable of running a multi-pass workflow to autonomously detect and correct its own errors. Cohere evaluates scores between 80 and 100 as either perfect or containing only minor errors, using GPT-5.6-Sol as the judge, though independent WMT26 results are still pending.
Mixture of Experts (MoE) architectures allow massive models to route tokens to specific expert subnetworks, activating only a fraction of the total parameters during inference. This design achieves the high accuracy of a giant model while maintaining the computational efficiency and lower operational costs of a smaller model.
Regionally, the standard model outperforms Gemma 4 31B across Europe, scoring 82.17 compared to Gemma's 72.73, while scoring closely in South Asia at 86.16 against Gemma's 88.04. Performance tests also show output speeds of 112 tokens per second at low concurrency compared to Gemma's 81, and 39 against 30 at high concurrency, delivering up to 1.4x higher throughput.
Long-document processing is another strength, with the model scoring 48.9 when translating two book chapters in a single API call, outperforming Google Translate at 21.3 and Gemma 4 31B at 19.4, as measured per paragraph using xCOMET-XL. Cost efficiency is also notable, averaging $0.000676 per task for roughly 661 tokens—making it about 58 times cheaper than Gemini 3.1 Pro Preview.
"Organizations that cannot communicate globally cannot stay sovereign."
Cohere
Developers can access North Small Translate instantly through Cohere's Chat V2 API, where it is available for free until rate limits are reached. The model is also available for non-commercial self-hosting and commercial licensing.
Source: MarkTechPost
Found something wrong in this article? Report an issue with this article
Comments
Leave a Comment