If your local coding agent prints a confident summary after a task, check your git diff. A new benchmark by developer Giuseppe Sirigu reveals that major agent frameworks, including the Linux Foundation’s Goose and Nous Research’s Hermes, often execute zero tool calls when paired with popular local models like qwen2.5-coder. The agents don’t crash; they simply narrate what they think they did while leaving your files completely untouched.
The Silent Failure Mode
Sirigu’s testing exposed a critical disconnect between model intent and runtime execution. When using qwen2.5-coder:7b, the Polyglot agent—which parses tool calls from raw text—achieved a 39% success rate. In stark contrast, pi, Hermes, and Goose all hit 0% task completion across 18 runs. The root cause? Hosted models like Claude use structured channels for tool calls, but open-weight models often emit JSON-like syntax as plain text. Runtimes listening only for native function-calling signals ignore this text, treating the model’s internal monologue as a final answer.
Scale Does Not Fix the Gap
You might assume this is a limitation of small models, but the failure persists even at 32B parameters. When Sirigu scaled the test to qwen2.5-coder:32b, Polyglot’s success rate jumped to 83%, proving the model can call tools. However, pi, Goose, Hermes, and opencode remained stuck at 0% completion. This indicates that raw model size does not automatically resolve the mismatch between how open-weight models emit instructions and how standard agent runtimes parse them.
Goose’s Shim and Hermes’ Parsers
Goose, backed by platinum members like AWS and Google, ships a specific fix called GOOSE_TOOLSHIM to handle non-tool-calling models. Even with this shim enabled, Goose managed only 1/18 tasks on the 14B model, often failing on schema mismatches. Hermes advertises "11 tool-call parsers," but Sirigu’s source code review shows these operate at the serving layer. The live agent loop actively deletes text-emitted tool calls as noise, a deliberate design choice that renders it blind to the very output local models produce.
Newer Models and Uneven Reliability
The story shifts with newer generations. On qwen3-coder (2025), native function-calling works well, with Polyglot and pi both achieving high success rates. The current hype-cycle model, Qwen3.8-27B, saw all agents hit 18/18 in Sirigu’s final tests. Yet, this reliability is uneven. Goose still scored 16/18 and Hermes 17/18 on the same model, revealing that variance is invisible until measured. The takeaway isn’t that every model is broken, but that you cannot assume your agent is working without verifying the actual file changes.
Key Takeaways
- Open-weight local models like qwen2.5-coder often emit tool calls as plain text, causing runtimes that rely on native function-calling channels to ignore them entirely.
- Increasing model size from 7B to 32B does not fix the tool-calling gap for agents like Goose and Hermes, which remained at 0% completion in the benchmark.
- Even with specific fixes like GOOSE_TOOLSHIM, major frameworks struggle to reliably execute tasks on older local model generations.
- Reliability varies significantly across model versions; while Qwen3.8-27B showed high success rates, minor variations in agent configurations led to inconsistent results.
The Bottom Line
Stop trusting your local agent's self-reported success. If you aren't parsing raw text for tool calls or verifying file changes directly, you are likely just reading a hallucination.
Sources
https://dev.to/gsirigu/your-local-coding-agent-might-be-doing-nothing-at-all-2hei