The most dangerous performance killers in local LLM inference are the ones that look like they’re working. A developer reported a 27B model running at Q4 on an RTX 3090, but throughput had collapsed from a healthy 60-70 tokens per second down to a crawl of 12-18. The configuration menu still screamed "GPU Offload: Max," yet the model was effectively running a hybrid CPU/GPU split. The culprit wasn’t a broken driver or a corrupted model file; it was LM Studio’s internal VRAM estimator quietly deciding that "Max" was too aggressive and trimming offload layers without warning.
The Log Tells the Truth, The UI Lies
Diagnostics are non-negotiable when your throughput drops off a cliff. The source of truth isn’t the dropdown menu; it’s the per-load server log located in the .lmstudio user profile directory. In this specific failure case, the logs revealed the estimator had adjusted the offload count from "max" down to "54" out of the model’s 65 layers. The line offloaded 54/65 layers to GPU is the smoking gun. When the requested offload doesn't match the actual offload, you are paying the CPU tax on every single token generated, resulting in a 2-3x slowdown that feels like a software regression but is actually a configuration mismatch.
CUDA 12 Runtime Over-Reserves Memory
The investigation pinpointed the CUDA 12 runtime backend as the trigger. The estimator appears to over-calculate memory usage for graph and KV buffers on this specific build, forcing the loader to be pessimistic and leave layers on the CPU to avoid an out-of-memory crash. Switching the runtime backend to plain CUDA or Vulkan restored full offload in every tested scenario. This simple dropdown change bumped throughput from roughly 15 tok/s back up to the 41.9-45.4 tok/s range across four different 27B models, with GPU utilization pinning near 97%.
Workarounds and The Open Issue
As of late September 2026, this remains an open issue in the LM Studio bug tracker with no official patch. Until a fix lands, users must rely on manual intervention. Setting an explicit layer count (e.g., forcing 65) can sometimes override the estimator, but if the log still reports a lower number, the runtime backend is still the bottleneck. Reducing context length helps free up VRAM for the weights, but the reporter noted it wasn't a reliable fix. The CLI command lms load is invaluable here, letting you see what the loader plans to do before you commit to a load.
Key Takeaways
- "Max" offload is a request, not a receipt: Always verify the
offloaded X/Y layersline in the logs. - Backend matters more than sliders: Switching from CUDA 12 to Vulkan/Plain CUDA resolved the 3x speed loss.
- Silent degradation is real: The app doesn't warn you when it trims layers to fit its own safety estimate.
- CLI diagnostics: Use
--estimate-onlyto debug offload plans without waiting for a full model load.
The Bottom Line
If you aren't reading the load logs, you are flying blind on your own hardware. The gap between a 60 tok/s model and a 15 tok/s model is often just a single runtime setting that the GUI refuses to explain.
Technical Context
This behavior highlights a broader issue in local LLM tooling: the conflict between safety and performance. The VRAM cap exists to prevent crashes, but when the estimator is wrong—as it appears to be with the CUDA 12 build—it silently sabotages performance. Users with AMD hardware face similar headaches, where ROCm might fail entirely while Vulkan works fine. The lesson is clear: bisect your runtime backends before assuming your model or hardware is at fault.