The prevailing narrative that small open-weight language models are too dim for agentic database work is facing a harsh reality check. A new paper from Cevheri Bozoglan, submitted to arXiv on September 18, 2026, suggests the bottleneck isn't the model's reasoning capacity, but the infrastructure wrapping it. By driving the agent mode of an open-source SQL client with 39 different locally served models over eleven days, the study uncovered that the majority of failures are mechanical, not cognitive.
The Myth of Capability Gaps
The dataset is massive: 8,199 runs, 110,711 ledger events, and 14,008 refused tool calls. Of the 2,100 losses attributed to specific models, 1,590 of themβ75.7%βoccurred in runs that had successfully invoked at least one tool. This statistic is critical. It implies the model knew what to do and tried to do it, but the execution environment failed. Only 17.3% of failures were classified as 'capability' issues, where the model invoked no tool at all. The authors caution that this specific ordering holds in only 74.5% of clustered resamples, but the dominance of non-capability failures is robust.
Invisible Server Defects
The largest failure class, 'transport' (36.2%), involves runs that used tools but never delivered results. Production ledgers typically record refusal codes but hide the model's actual arguments, masking the root cause for ten days. Once the researchers captured these arguments, they exposed five distinct server defects. One particularly egregious bug demanded a field on one tool, forbade it on the sibling tool that composed it, and then failed the run for its absence. Fixing these server issues alone moved six models by 6 to 21 cells out of 30 in performance rankings, touching no prompts or sampling settings.
The Memory Confound
Beyond logic errors, the study highlights a dangerous confound in local benchmarking: context window management. Without a hard cap, one 7.1 GB model was admitted at its full 262,144-token window, consuming 51 GB of RAM on a 64 GB machine. This created runs that were indistinguishable in logs from simple timeouts. For builders deploying small models locally, this is a wake-up call. Your agent isn't dumb; it's being crushed by its own context window or tripped by a server bug that doesn't know its own API schema.
Key Takeaways
- 75.7% of model-attributed failures occurred after a tool was invoked, pointing to infrastructure, not intelligence.
- Five server defects were identified by capturing previously hidden model arguments, including a contradictory API schema.
- Server-only fixes improved performance for six models by up to 21 cells without changing the model or prompt.
- Unbounded context windows can cause memory exhaustion (51 GB on a 64 GB machine) that mimics model timeouts.
- The corpus, scorer, and verifier are released for reproducibility, challenging the 'small models can't agent' dogma.
The Bottom Line
Stop blaming your 7B model for your agent's failure. You're likely fighting a broken server schema or a memory leak, not a reasoning deficit.