The repo Strata recently hit 788 points on Hacker News with a claim that sounds like a typo: running Qwen 3.8 Flash Next, a 125B-parameter model, on a consumer RTX 4090 at 100 tokens per second. While the throughput numbers are real, the headline obscures the critical infrastructure requirement: this performance is only achievable with a massive amount of system RAM, not just the GPU alone. The claim separates the physics of VRAM limits from the marketing of "consumer hardware."
The Sparse MoE Architecture Trick
Qwen 3.8 Flash Next is a sparse Mixture-of-Experts (MoE) model with approximately 180 billion total parameters, including 125B for the main model, 51B for n-gram embeddings, and 4B for the multi-token prediction head. However, only about 6 billion parameters are active per token. Strata exploits this by keeping hot weights, such as attention and shared experts, in VRAM while parking the rest in system RAM and streaming them over PCIe. The n-gram table, which accounts for a third of the parameters, functions as a lookup rather than a matrix multiplication, allowing it to reside on NVMe storage without significantly degrading throughput.
The Hidden Cost of System RAM
The title's omission of RAM requirements is the most significant caveat. Strataβs documentation specifies that a 32 GB RAM setup is only suitable for a "Coder" variant with half the experts removed. For the full IQ3_S quantization recommended for balance, 64 GB of system RAM is required, while the highest quality UD-IQ4_XS needs 96 GB or more. An independent benchmark of the RTX 4090 setup utilized a Ryzen 7900 with 192 GB of DDR5 RAM, power-capped at 280 W, to achieve 110.8 tok/s on text decode. This confirms that the bottleneck has shifted from VRAM capacity to system memory bandwidth and capacity.
Workload-Dependent Performance
The advertised 100 tok/s is not a universal constant; it is heavily dependent on the workload and speculative decoding acceptance rates. On a Strix Halo box, the model achieved roughly 22 tok/s on prose but jumped to 82 tok/s on code and 93 tok/s on structured output. This variance occurs because multi-token prediction and speculative decoding have higher acceptance rates on predictable formats. Users should expect significantly slower speeds for chatty prose compared to code generation, making the headline a best-case scenario rather than an average performance metric.
Key Takeaways
- Strata runs Qwen 3.8 Flash Next by offloading inactive experts and n-gram tables to system RAM and NVMe.
- Achieving ~100 tok/s requires at least 64 GB of system RAM for the recommended IQ3_S quant.
- Performance varies drastically by workload: 22 tok/s on prose vs. 82+ tok/s on code.
- The model uses the Qwen Community License 1.0, not Apache 2.0, with restrictions for large commercial products.
The Bottom Line
This isn't a miracle of GPU efficiency; it is a triumph of memory hierarchy engineering that redefines "local AI" to include heavy system RAM requirements. If you don't have 64 GB+ of DDR5, this benchmark is irrelevant to your hardware.