Running a 510 GB mixture-of-experts model on an 8 GB GPU sounds impossible, but one developer proved it viable on an RTX 5060 using a Gen5 NVMe drive and 31 GiB of system RAM. The project, documented by helgard_orlm, achieves roughly 2.4 tokens per second in live chat by streaming expert weights directly from disk and caching frequently used experts in RAM. The primary challenge was not just speed, but stability: the hardware is so new that standard inference kernels fail silently, producing garbage output without throwing errors.
The Physics of Slow Inference
The bottleneck is I/O bandwidth rather than compute power. DeepSeek-V4.1-Flash requires pulling 4.2 GiB of data per generated token because it activates 6 experts across 40 layers, with each expert weighing 17.93 MiB. With the drive delivering 8.2 GiB/s, the physical act of reading data takes about 0.5 seconds per token, leaving minimal room for software optimization. The team optimized this by overlapping disk reads with GPU computation, reducing wait times from 0.65 to 0.31 seconds per token.
Three Silent Bugs on Blackwell Hardware
The most valuable contribution of this project is the documentation of three specific failures on RTX 50xx (sm_120) cards. First, TileLang 0.1.8, the pinned version in DeepSeek's reference code, produced cosine similarities of 0.0006 against PyTorch, effectively generating random noise instead of math. Second, a race condition in the act_quant kernel caused NaN values in FP8 outputs for batches larger than 64 rows, a bug that only appeared during prompt processing, not single-token generation. Third, the sparse attention kernel requested 141 KB of shared memory, exceeding the ~99 KB available on consumer Blackwell cards, requiring a workaround to process heads in smaller groups.
Collaboration Between Human and Agents
This build was a hybrid workflow involving Claude (Anthropic) and Codex (OpenAI). The developer set the goals and model choice, while Claude handled the engine, server, and bug hunts. Codex contributed the critical overlap scheme and the one-token short path optimization. The team verified that every optimization kept the output token-for-token identical to the baseline, ensuring that speed gains did not come at the cost of model fidelity.
Key Takeaways
- Check bytes per token, not just parameter count: The Qwen model (1.27 GiB/token) runs 3.3x faster on the same hardware than DeepSeek-V4.1-Flash (4.2 GiB/token).
- Trust no kernel on new hardware: Always compare custom kernels against a plain PyTorch implementation on real weights, as self-tests often check shapes rather than numerical values.
- Overlap is critical: Using CUDA events to overlap disk reads with computation increased throughput from 0.89 to 1.28 tokens per second in the first optimization step.
- Silent failures are the norm: On sm_120 architectures, incorrect math often presents as nonsensical text or NaNs rather than explicit exceptions.
The Bottom Line
If you are building a local AI rig, stop looking at parameter counts and start measuring bytes per token; for MoE models, I/O bandwidth is the only metric that matters. Furthermore, treat new GPU architectures like Blackwell as experimental terrain where silent numerical corruption is the default state until proven otherwise by rigorous torch-based validation.