Skip to main content

Meta Releases Llama 4 Scout Open-Weights 105B Model

Meta FAIR introduces Llama 4 Scout, a 105B open-weights model featuring dynamic Mixture-of-Depths architecture and native 1M token context support.

AI-written
Inewgen
11 Oct 2026Source: Dev.to4 min read (0 views)
Share
Meta Releases Llama 4 Scout Open-Weights 105B Model

Stock photo for illustration only, not from the actual event

Font size
  • Meta FAIR launches Llama 4 Scout as the open community's first MoD foundation model
  • Features 105 billion total parameters with only 24 billion active parameters per token
  • Supports a native 1 million token context window with zero retrieval degradation
  • Requires only 24GB VRAM to run locally on a single consumer GPU workstation

Meta's Fundamental AI Research (FAIR) team has officially rolled out Llama 4 Scout, marking the open-source community's initial frontier-class foundation model engineered natively on dynamic Mixture-of-Depths or MoD routing. Boasting a total parameter capacity of 105 billion alongside just 24 billion active parameters per token pass, Llama 4 Scout achieves complete parity with leading closed-source reasoning engines across synthetic coding, advanced mathematics, and long-context synthesis benchmarks while executing efficiently on enterprise dual-GPU setups.

Standard dense transformer architectures allocate an identical compute budget to every single token, evaluating routine punctuation and filler words with the exact same 80-layer depth assigned to complex mathematical proofs. Llama 4 Scout discards this uniform resource distribution by embedding a learned top-k routing gate directly into each transformer block.

artificial intelligence neural network chip technology

Stock photo for illustration only, not from the actual event

During runtime inference, the internal router evaluates token complexity and dynamically assigns computational resources. Simple tokens bypass deep self-attention and feed-forward layers via residual shortcut paths, whereas highly informative tokens secure maximal compute across specialized reasoning blocks. This dynamic depth allocation reduces generation latency by 58% compared to standard dense 70B models while expanding total model capacity to 105 billion parameters.

105BTotal Parameters
24BActive Parameters / Token
1MNative Context Window

Regarding architectural mechanics, the model integrates RingAttention coupled with sliding-window multi-head latent attention or MLA. Official FP8 and INT4 AWQ quantization checkpoints are directly accessible for download via Hugging Face, requiring a mere 24GB VRAM inference footprint in 4-bit mode for single GPU execution.

advanced computer server rack hardware workspace

Stock photo for illustration only, not from the actual event

Addressing the challenge of maintaining native 1M-token context windows without performance loss, previous open-weights models claiming multi-hundred-thousand token support often relied on post-training RoPE interpolation techniques like YaRN, which triggered sharp perplexity spikes beyond 64K tokens. Conversely, Llama 4 Scout underwent pre-training from step zero utilizing an 8-stage progressive context expansion curriculum scaling up to 1,048,576 tokens.

Never miss the latest news?

Subscribe to get news summaries by email - not often enough to be annoying.

โฆษณา

The integration of dynamic Mixture-of-Depths routing combined with an end-to-end trained 1-million token context window represents a major paradigm shift for open-weights AI. By decoupling parameter capacity from active per-token compute, Meta enables smaller teams and data-sensitive enterprises to deploy frontier-grade intelligence locally without incurring high recurring API costs.

On the industry-standard Needle In A Haystack benchmark spanning the entire 1M token spectrum, Llama 4 Scout registered a 99.8% retrieval accuracy across 2,500 distinct document depths, effortlessly processing extensive legal corpora, complete software repositories, and massive scientific literature collections. Furthermore, independent third-party evaluations on SWE-bench Verified showed the model successfully resolving 48.6% of real-world GitHub issues autonomously.

Base, instruction-tuned, and quantized checkpoints are now globally accessible on the Hugging Face Hub and Meta's official distribution portal, with native runtime support already merged into vLLM, Ollama, LM Studio, and Hugging Face TGI for immediate enterprise Kubernetes deployment.

Source: Dev.to

Comments

Leave a Comment
0/2000

Found something wrong in this article? Report an issue with this article