Building large-scale LLM applications often hits a hard wall: the cost of recomputing key-value (KV) states for massive context windows. A new paper submitted to arXiv on October 7, 2026, by Sietse Schelpe introduces galahad-kv, a public package that bypasses this bottleneck by storing KV states on encrypted local NVMe disk. The result is a memory layer that loads saved states byte-exact, eliminating the need for expensive GPU recomputation.

Breaking The Recompute Bottleneck

The core innovation here is practical infrastructure optimization. Instead of forcing the GPU to recalculate internal states every time a prompt is sent, galahad-kv saves the KV state of each ~16,000-token block. In tests running on a single NVIDIA H100 with Gemma 4 12B and 31B models, the system processed a 50,000,000-token stream. Every block probed from depths ranging from 0 to 50M tokens was retrieved successfully, with a 100% success rate (100 of 100) across both models.

Efficiency Gains And Accuracy Metrics

For developers watching their cloud bills, the efficiency numbers are compelling. Loading a block from the encrypted store was 2.8x to 4.3x faster than recomputing it, while consuming 8.8x to 12.3x less GPU energy. Critically, GPU memory usage remained flat throughout the entire 50M-token stream, solving the linear memory scaling problem that plagues long-context inference. On the accuracy front, when asked about facts planted millions of tokens earlier, the Gemma 4 31B model answered correctly 98 times out of 100, while the 12B model achieved 82 correct answers. Neither model hallucinated a response.

Key Takeaways

  • galahad-kv reduces GPU energy consumption by up to 12.3x compared to standard recomputation methods.
  • The system supports 50M-token contexts on a single NVIDIA H100 GPU without memory blowout.
  • Retrieval latency is improved by 2.8x to 4.3x, making long-context queries significantly faster.
  • The package is open-source under a free license and uses public software like vLLM for reproduction.
  • Accuracy remains high, with the 31B model achieving a 98% success rate on long-distance fact retrieval.

The Bottom Line

This isn't magic; it's smart infrastructure engineering. By offloading KV states to cheap, fast NVMe storage, we can finally scale long-context AI without melting our GPUs or bankrupting our budgets. It's a pragmatic win for builders who need persistent, efficient memory in their LLM stacks.