Skip to main content

I Built a 188M Mixture-of-Experts LLM on a Free GPU

Shivam Kumar, founder of VisionQuantech, built a sparse 188M Mixture-of-Experts LLM from scratch on a free Tesla T4 GPU for zero dollars.

AI-written
Inewgen
07 Oct 2026Source: Dev.to2 min read (0 views)
Share
I Built a 188M Mixture-of-Experts LLM on a Free GPU

Stock photo for illustration only, not from the actual event

Font size
  • Shivam Kumar built a 188M parameter MoE LLM on a free Tesla T4 GPU
  • Used the Main Researcher System v4 methodology for systematic design
  • Employed 64 micro-experts to maximize capacity at a low compute cost
  • Trained on approximately 215 million tokens with a zero-dollar budget

Cutting-edge Mixture-of-Experts (MoE) research typically demands massive billion-parameter scales, keeping it far out of reach for a single free Colab GPU. However, the core principle of MoE remains elegant: only a fraction of the model parameters need to activate for each individual token processed.

Driven by this approach, Shivam Kumar, founder of VisionQuantech, set out to build a sparse 188M parameter MoE model entirely from scratch using a free Tesla T4 GPU with a zero-dollar budget, aiming to prove that tiny architectures can be engineered rigorously rather than purely through trial and error.

server room data center no logo

Stock photo for illustration only, not from the actual event

To design the architecture, Kumar utilized a methodology termed the Main Researcher System v4, combining 12 MoE primitives extracted from existing literature with 11 meta-patterns, including TRIZ contradiction resolution and morphological analysis, to resolve the trade-offs between model capacity and routing overhead.

Deploying an MoE architecture on constrained consumer-grade hardware highlights a creative approach to resource efficiency. By breaking down coarse experts into numerous micro-experts, the model can capture complex linguistic patterns and syntax without incurring the heavy compute penalties typically required by dense models.

Initial CPU tests spanning 300 steps showed that all candidate models trained stably, with loss dropping from 41.5 down to 7.6 without encountering NaN values or model collapse. The winning configuration achieved an aux loss of 2.03 and left zero experts starving of utilization.

Never miss the latest news?

Subscribe to get news summaries by email - not often enough to be annoying.

โฆษณา

"This is the honest version — what's proven, what's measured, and what's still running."

Shivam Kumar

The active training pipeline on the free Colab T4 consists of two main phases: Phase 1 utilizes TinyStories with roughly 200 million tokens for linguistic fluency, followed by Phase 2 using 15 million instruction tokens, summing up to roughly 215 million total tokens over a 2.5 to 3-hour run.

Source: Dev.to

Comments

Leave a Comment
0/2000

Found something wrong in this article? Report an issue with this article