Skip to main content

Cohere Releases North Small Translate 218B MoE Model

Cohere introduces North Small Translate, a 218B MoE translation model scoring 83.6 on WMT26 across 50 languages with an Agentic variant.

AI-written
Inewgen
12 Sep 2026Source: MarkTechPost3 min read (0 views)
Share
Cohere Releases North Small Translate 218B MoE Model

Stock photo for illustration only, not from the actual event

Font size
  • North Small Translate is the first translation model in Cohere's North product family.
  • It uses a decoder-only sparse MoE architecture with 218B total and 25B active parameters.
  • The model achieves an average score of 83.6 on WMT26 benchmarks across 50 languages.

Cohere has officially released North Small Translate, a decoder-only sparse Mixture of Experts (MoE) translation model featuring a total of 218 billion parameters. Approximately 11.5% of its weights are active per token, meaning per-token compute tracks the 25 billion active parameters while its memory footprint requires holding all 218B weights.

This model marks the first translation release within Cohere's North family, succeeding earlier multilingual lineages such as Tiny Aya and Command A Translate. Cohere partnered with RWS and its Language Weaver scientists and language experts to refine the model's real-world translation quality, emphasizing that global communication is foundational to organizational sovereignty.

software engineer programming code screen workspace

Stock photo for illustration only, not from the actual event

83.6WMT26 All-Languages Score
11.5%Active Weights Per Token
1.4xMax Throughput Multiplier

In terms of performance, the model reports a WMT26 all-languages score of 83.6 across 50 supported languages. It also features an Agentic variant capable of running a multi-pass workflow to autonomously detect and correct its own errors. Cohere evaluates scores between 80 and 100 as either perfect or containing only minor errors, using GPT-5.6-Sol as the judge, though independent WMT26 results are still pending.

Mixture of Experts (MoE) architectures allow massive models to route tokens to specific expert subnetworks, activating only a fraction of the total parameters during inference. This design achieves the high accuracy of a giant model while maintaining the computational efficiency and lower operational costs of a smaller model.

Never miss the latest news?

Subscribe to get news summaries by email - not often enough to be annoying.

โฆษณา

Regionally, the standard model outperforms Gemma 4 31B across Europe, scoring 82.17 compared to Gemma's 72.73, while scoring closely in South Asia at 86.16 against Gemma's 88.04. Performance tests also show output speeds of 112 tokens per second at low concurrency compared to Gemma's 81, and 39 against 30 at high concurrency, delivering up to 1.4x higher throughput.

Long-document processing is another strength, with the model scoring 48.9 when translating two book chapters in a single API call, outperforming Google Translate at 21.3 and Gemma 4 31B at 19.4, as measured per paragraph using xCOMET-XL. Cost efficiency is also notable, averaging $0.000676 per task for roughly 661 tokens—making it about 58 times cheaper than Gemini 3.1 Pro Preview.

"Organizations that cannot communicate globally cannot stay sovereign."

Cohere

Developers can access North Small Translate instantly through Cohere's Chat V2 API, where it is available for free until rate limits are reached. The model is also available for non-commercial self-hosting and commercial licensing.

Source: MarkTechPost

Comments

Leave a Comment
0/2000

Found something wrong in this article? Report an issue with this article