Skip to main content

Darwin-180B-RSI: VIDRAFT's MoE Model Claims Five Tops

Korean Pre-AGI startup VIDRAFT has released Darwin-180B-RSI, a 180B MoE reasoning model featuring recursive self-improvement and five Hugging Face leaderboard leads.

AI-written
Inewgen
30 Sep 2026Source: Dev.to3 min read (0 views)
Share
Darwin-180B-RSI: VIDRAFT's MoE Model Claims Five Tops

Stock photo for illustration only, not from the actual event

Font size
  • VIDRAFT, a South Korean Pre-AGI startup, released Darwin-180B-RSI in late September 2026
  • The model features 180 billion parameters with a sparse activation of 10 out of 512 experts per token
  • It reportedly claims the number one spot on five Hugging Face leaderboards across math and science
  • Combines selective model merging and recursive self-improvement through verifiable correctness

Artificial intelligence development has reached another milestone as VIDRAFT, a South Korean Pre-AGI startup, introduced Darwin-180B-RSI, an open reasoning model built on the Qwen3.8-Flash-Next architecture. The model combines a massive 180-billion-parameter scale with a unique synthesis of selective model merging and recursive self-improvement (RSI).

In terms of structural design, Darwin-180B-RSI implements a sparse expert activation framework. By routing each token through only 10 out of 512 total experts, the effective compute required per forward pass remains significantly lower than a traditional dense 180-billion-parameter model, offering practical advantages in deployment costs and latency.

artificial intelligence neural network data visualization

Stock photo for illustration only, not from the actual event

180BTotal Parameters
10/512Active Experts per Token
5Hugging Face Leaderboard Tops

Conceptually, the pipeline merges selective model merging and recursive self-improvement, aligning closely with ideas like rejection sampling fine-tuning or self-play reinforcement learning. However, specific implementation details such as iteration counts, reward modeling parameters, and filtering criteria were not disclosed in the initial reporting.

Never miss the latest news?

Subscribe to get news summaries by email - not often enough to be annoying.

โฆษณา

Mixture of Experts (MoE) architectures allow massive parameter scaling while keeping inference compute tractable per token. Nevertheless, engineers must still ensure sufficient VRAM capacity to hold all expert weights in memory, making quantized deployment strategies or expert offloading essential considerations for self-hosting.

Despite claiming top positions across five Hugging Face leaderboards spanning mathematics, science, and general reasoning, the source report highlights an important caveat: these benchmark scores are self-reported by VIDRAFT without independent third-party validation. Performance on static benchmarks may not generalize seamlessly to proprietary production workloads or out-of-distribution tasks.

The model is designated as an open model available on Hugging Face. Developers can check repository availability and download weights using standard toolchain commands such as huggingface-cli search vidraft/darwin-180b-rsi or huggingface-cli download vidraft/darwin-180b-rsi, though exact repository paths should be verified directly on Hugging Face.

Source: Dev.to

Comments

Leave a Comment
0/2000

Found something wrong in this article? Report an issue with this article