Running AI agents on consumer hardware is less about raw power and more about brutal arithmetic. A recent homelab census by user c1-anderson highlights the stark reality of local LLMs: a 27-billion-parameter model loaded with a total footprint of 18.3 GB had only 0.5 GB residing on the GPU. The remaining 17.8 GB sat in system RAM, forcing the CPU to grind through tokens one by one. This isn't running a model; it's running a bottleneck.
The Hardware Reality Check
The setup comprises four boxes with 26 CPU threads and 71 GB of RAM, hosting 41 Docker containers and 9 LXC containers. The core AI workload runs on an LLM box equipped with an i7-10875H, 31 GB of RAM, and an RTX 2060 with just 6 GB of VRAM. The census, measured on September 24, revealed that the agent framework requires a minimum 64K context window, a requirement that clashes violently with the limited VRAM. The best local model, qwen3.5 9B, only achieves 51% GPU utilization at 64K context, spilling the rest to the CPU.
Routing Decisions and Cloud Fallbacks
The performance gap between local and cloud inference dictates the architecture. A typical agent turn on the local 9B model took roughly 90 seconds, whereas a cheap cloud model via OpenRouter answered tool-call probes in 1.5 to 2 seconds. Consequently, the general agent profiles (chief of staff, research, writing) now route to OpenRouter's qwen3.7-flash, costing approximately $1.39 a month based on 30 days of usage. The coding agent moved to an OpenCode Go subscription for $10 a month to avoid metered bills, while security audits remain strictly local to prevent data egress.
The Silence Failure and Watchdog Logic
Local AI is prone to silent failures rather than explicit errors. A pinned model can hang Ollama without logging anything, causing all other requests to queue indefinitely. The author implemented a script to unload models pinned for more than 48 hours to mitigate this. Additionally, a watchdog system failed on September 7 when an upstream PII-redaction filter caused health checks to fail across all paid models. This incorrectly routed four agents back to the slow local model for 17 days until manual intervention, highlighting the fragility of automated failover logic.
Key Takeaways
- VRAM is the primary constraint: 6 GB is insufficient for full GPU offloading of 27B models, making them effectively CPU-bound.
- Input token cost dominates: Agent workloads are prompt-heavy, making input pricing the key metric for cloud model selection.
- Silent hangs are dangerous: Unlike API errors, local model hangs do not trigger fallbacks without active monitoring scripts.
- Cloud is faster, local is private: The architecture correctly splits workloads based on latency requirements and privacy needs.
The Bottom Line
Local AI on consumer hardware is a privacy compromise, not a performance one. If you cannot fit your model's active context window into VRAM, you are building a bottleneck, not an agent.