The choice between Ollama, vLLM, and llama.cpp is no longer a simple matter of which engine is "faster." A new benchmark published on DEV.to on October 1, 2026, demonstrates that the primary differentiator is concurrency handling, not raw inference speed. For single-user local setups, the performance delta is negligible and driven almost entirely by default configuration settings. However, for multi-user environments, the architectural differences in batch processing and memory management create massive disparities in throughput and latency.
The Benchmark: Defaults Over Engines
The test compared Ollama 0.34.4 and llama.cpp 0.5.0 on an Apple M1 with 16 GB of RAM, using the same Qwen3-4B GGUF file quantized to Q4_K_M. In a single-request scenario, Ollama outperformed llama.cpp with 18.7 tokens per second versus 13.7 tokens per second. This gap was not due to superior engine code, but rather different default context lengths: Ollama auto-selected a 4,096-token context, while llama.cpp configured four parallel slots with a 40,960-token context each. The results highlight that comparing these tools without normalizing for context size and parallelism settings is technically flawed.
Concurrency Is the Real Battleground
The performance hierarchy flips dramatically under load. When four concurrent requests were sent, default Ollama (with OLLAMA_NUM_PARALLEL=1) queued the requests, resulting in an aggregate throughput of 18.9 tokens per second and a 54-second wall time. In contrast, llama.cpp, which defaults to auto-sized parallel slots, achieved 28.9 tokens per second and finished in 35 seconds. However, tuning Ollama with OLLAMA_NUM_PARALLEL=4 yielded the best result: 35.0 tokens per second aggregate and a 29-second completion time. This proves that with minor configuration tweaks, the "slower" engine can dominate the "faster" one.
vLLMβs Aggressive Memory Strategy
vLLM operates on a different paradigm entirely, optimized for high-throughput serving on GPU servers rather than local laptops. Its headline features, including PagedAttention and continuous batching, allow it to serve many concurrent users efficiently. However, vLLMβs default --gpu-memory-utilization flag is set to 0.92, meaning it claims 92% of available GPU memory for the model and KV cache. This aggressive allocation ensures high concurrency but makes running multiple vLLM instances or other workloads on the same GPU difficult without manual memory splitting. Furthermore, vLLMβs support for GGUF models remains labeled as "highly experimental and under-optimized," requiring a separate plugin package.
Key Takeaways
- Single-user performance between Ollama and llama.cpp is determined by default context and slot settings, not engine architecture.
- Ollamaβs default OLLAMA_NUM_PARALLEL=1 creates a bottleneck for concurrent users; raising this value significantly improves aggregate throughput.
- vLLM is the only viable option for high-concurrency GPU serving but claims 92% of GPU memory by default and has poor GGUF support.
- Apple Silicon users should stick to Ollama or llama.cpp, as vLLMβs Apple Silicon support is experimental and CPU-bound.
The Bottom Line
Stop benchmarking local LLM engines on idle machines with single requests; that data is noise. The real test is how each engine handles concurrent connections and memory allocation. For most developers, the choice is between Ollamaβs convenience with manual tuning and vLLMβs raw power on dedicated GPUs. llama.cpp remains the king of control, but its defaults are often misconfigured for modern multi-user expectations.