I Built a 188M Mixture-of-Experts LLM on a Free GPU
Shivam Kumar, founder of VisionQuantech, built a sparse 188M Mixture-of-Experts LLM from scratch on a free Tesla T4 GPU for zero dollars.

Stock photo for illustration only, not from the actual event
- Shivam Kumar built a 188M parameter MoE LLM on a free Tesla T4 GPU
- Used the Main Researcher System v4 methodology for systematic design
- Employed 64 micro-experts to maximize capacity at a low compute cost
- Trained on approximately 215 million tokens with a zero-dollar budget
Cutting-edge Mixture-of-Experts (MoE) research typically demands massive billion-parameter scales, keeping it far out of reach for a single free Colab GPU. However, the core principle of MoE remains elegant: only a fraction of the model parameters need to activate for each individual token processed.
Driven by this approach, Shivam Kumar, founder of VisionQuantech, set out to build a sparse 188M parameter MoE model entirely from scratch using a free Tesla T4 GPU with a zero-dollar budget, aiming to prove that tiny architectures can be engineered rigorously rather than purely through trial and error.

Stock photo for illustration only, not from the actual event
To design the architecture, Kumar utilized a methodology termed the Main Researcher System v4, combining 12 MoE primitives extracted from existing literature with 11 meta-patterns, including TRIZ contradiction resolution and morphological analysis, to resolve the trade-offs between model capacity and routing overhead.
Deploying an MoE architecture on constrained consumer-grade hardware highlights a creative approach to resource efficiency. By breaking down coarse experts into numerous micro-experts, the model can capture complex linguistic patterns and syntax without incurring the heavy compute penalties typically required by dense models.
Initial CPU tests spanning 300 steps showed that all candidate models trained stably, with loss dropping from 41.5 down to 7.6 without encountering NaN values or model collapse. The winning configuration achieved an aux loss of 2.03 and left zero experts starving of utilization.
"This is the honest version — what's proven, what's measured, and what's still running."
Shivam Kumar
The active training pipeline on the free Colab T4 consists of two main phases: Phase 1 utilizes TinyStories with roughly 200 million tokens for linguistic fluency, followed by Phase 2 using 15 million instruction tokens, summing up to roughly 215 million total tokens over a 2.5 to 3-hour run.
Source: Dev.to
Found something wrong in this article? Report an issue with this article
Comments
Leave a Comment