Aleph Alpha Releases Kolibri: 78.1B Open-Weight MoE Model
Aleph Alpha has launched Kolibri, a 78.1B Mixture-of-Experts English-German model activating only 3.46B parameters per token, running on a single B200 or H200 under Apache 2.0.

Stock photo for illustration only, not from the actual event
- Aleph Alpha released Kolibri, a 78.1B-parameter Mixture-of-Experts model
- Activates only 3.46B parameters per token for high efficiency
- Supports a massive 1-million token context window with per-request reasoning
- Available under Apache 2.0 with FP8 weights, runnable on a single B200 or H200
AI firm Aleph Alpha has officially announced the release of Kolibri, a cutting-edge English-German Mixture-of-Experts (MoE) language model boasting a total of 78.1 billion parameters. The model is built to bridge advanced bilingual capabilities with optimized computational efficiency for enterprise and developer use.
The defining characteristic of Kolibri lies in its sparse activation architecture. Despite its massive total size, the model activates a mere 3.46 billion parameters per token during inference. This selective activation drastically lowers compute overhead, ensuring rapid response times while preserving the depth and breadth of a much larger network.
Beyond raw parameter efficiency, Kolibri introduces robust long-context handling, supporting up to 1 million tokens in a single request. It also incorporates per-request reasoning effort adjustments, allowing the system to scale its computational depth dynamically based on the complexity of the user query.

Stock photo for illustration only, not from the actual event
The deployment of Mixture-of-Experts models that activate only a fraction of their total parameters represents a crucial shift toward sustainable AI infrastructure. By enabling high-performance models to run on heavily optimized hardware configurations without massive cluster requirements, companies like Aleph Alpha are democratizing access to state-of-the-art language capabilities.
Kolibri's FP8 weights are released under the permissive Apache 2.0 license, allowing developers and organizations to inspect, modify, and deploy the model freely. Notably, the entire system is optimized to run efficiently on a single enterprise-grade AI accelerator such as the NVIDIA B200 or H200, significantly reducing hardware deployment barriers.
Source: MarkTechPost
Found something wrong in this article? Report an issue with this article
Comments
Leave a Comment