Skip to main content

VIDRAFT POCKET-Darwin-180B-GGUF: MoE Model on Laptops

VIDRAFT releases POCKET-Darwin-180B-GGUF, a 4-bit quantized 180B MoE model that runs on gaming laptops with 32GB RAM and 8GB VRAM.

AI-written
Inewgen
07 Oct 2026Source: Dev.to3 min read (0 views)
Share
VIDRAFT POCKET-Darwin-180B-GGUF: MoE Model on Laptops

Stock photo for illustration only, not from the actual event

Font size
  • VIDRAFT released POCKET-Darwin-180B-GGUF, a 180-billion-parameter MoE model
  • Quantized to 4-bit, reducing storage footprint from 360GB down to 111GB
  • Utilizes llama.cpp SSD streaming to load only required expert weights per token
  • Runs locally on consumer hardware like a gaming laptop with 32GB RAM and 8GB VRAM

The artificial intelligence community is buzzing following VIDRAFT's release of POCKET-Darwin-180B-GGUF, a compressed and quantized 4-bit GGUF version of their foundation model Darwin-180B-RSI, which was originally built upon the Qwen3.8-Flash-Next base. The original Darwin-180B-RSI weighs 360GB at BF16 precision and conventionally demands a multi-GPU server cluster to operate. The POCKET variant directly tackles these heavy infrastructure barriers, following VIDRAFT's product strategy of RSI for capability and POCKET for deployability.

The core breakthrough allowing such a massive model to function on local consumer hardware stems from combining three distinct techniques: a Sparse Mixture-of-Experts (MoE) architecture, 4-bit quantization, and llama.cpp-based SSD weight streaming. The full model comprises 512 expert sub-networks, but for every single token generated, a learned router activates only 10 of them. Consequently, the actual compute cost per token corresponds to approximately 3 billion active parameters rather than the full 180 billion total parameter count.

chromebook notebook computer office desk workspace

Stock photo for illustration only, not from the actual event

On the storage side, converting the model weights from BF16 to 4-bit integers slashes the footprint from 360GB to 111GB while preserving the MoE routing structure. Rather than loading the entire 111GB file into main memory at inference time, the system leverages llama.cpp's capability to stream weights dynamically from an SSD on demand. Only the specific expert weights required for the current token need to reside in memory at any given moment, effectively turning the SSD into an extension of the memory hierarchy to support the 8GB VRAM and 32GB RAM laptop configuration.

Never miss the latest news?

Subscribe to get news summaries by email - not often enough to be annoying.

โฆษณา

180BTotal Parameters
~3BActive Parameters per Token
111GBStorage Size after 4-bit

VIDRAFT published self-reported evaluation figures from their testing on the MMLU-Pro benchmark using 2,000 matched items, observing no measurable accuracy degradation under their evaluation methodology. Developers should note that independent replication is always encouraged for customized workloads. The model is publicly available on Hugging Face and ModelScope in GGUF format, requiring a recent llama.cpp build and roughly 111GB of free storage space to run via command-line tools.

Running a 180-billion-parameter model on consumer-grade hardware highlights a major shift toward efficient edge deployment. By leveraging sparse activation alongside demand-paged disk streaming, organizations requiring strict air-gapped security or local data processing can bypass heavy cloud dependencies without sacrificing model scale.

Developers interested in deployment should consult the official model card on Hugging Face for exact file names, recommended context length configurations, and specific chat templates tailored to their specific hardware setups.

Source: Dev.to

Comments

Leave a Comment
0/2000

Found something wrong in this article? Report an issue with this article