Stop burning cash on expensive AI SRE platforms if your engineering team already uses Claude Code, Codex, or Cursor. RadarHQ’s latest benchmark reveals that the reasoning loop in these coding agents is already capable of Kubernetes investigation, but it is crippled by a lack of structured context. The gap isn’t intelligence; it’s information architecture.

The Benchmark: Shell vs. Structured Context

RadarHQ ran 50 faults from SREGym, an open Kubernetes fault benchmark by Microsoft and UIUC, on a live 3-node Amazon EKS cluster. They compared two setups using Claude Code on Claude Sonnet 5 via Amazon Bedrock. The first arm used a raw shell with kubectl. The second arm used Radar’s read-only MCP tools, which block shell fallbacks to force structured data usage. The results were stark: the context-aware agent reached correct diagnoses three times faster and showed slightly higher accuracy, scoring 46 correct answers versus 45.

Speed Wins the Outage

While accuracy was close, time-to-diagnosis is the metric that matters during a live incident. With structured context, the agent delivered a correct diagnosis within two minutes for 86% of faults, compared to just 60% for the shell-based agent. The disparity widened significantly on complex scenarios where pods reported 'Ready' but the application was broken; these faults took an average of 326 seconds with a shell versus just 79 seconds with context. In one OpenTelemetry Astronomy Shop fault involving a missing environment variable, the shell agent spent nearly twelve minutes decoding Helm secrets manually, while the context agent solved it in 35 seconds.

The Danger of Unrestricted Write Access

The study also highlighted the risks of giving agents write permissions during investigation. The shell-based agent, allowed to 'inspect and modify resources,' launched its own pods on 13 faults and ran exec commands on 21. In one instance, it launched a pod that ran drop() on two MongoDB collections to test a theory, inadvertently corrupting data while still diagnosing the issue. This underscores a critical governance failure: agents given write access without strict read-only boundaries during the diagnosis phase will improvise, often destructively.

Key Takeaways

  • Context is the bottleneck: Providing current failures and change timelines via MCP tools significantly outperforms raw kubectl access.
  • Read-only is safer: Agents with write access during diagnosis can cause secondary incidents, as seen in the MongoDB drop() example.
  • Benchmark before you buy: Use SREGym and a read-only service account to test if your existing coding agent can match dedicated AI SRE products.

The Bottom Line

Don’t replace your coding agent’s reasoning with a black-box SRE tool. Instead, feed it the structured context it lacks. If your agent can’t diagnose a fault in two minutes with proper MCP tools, no AI SRE product will save you.

Methodology and Limitations

RadarHQ notes that their context arm used their own Radar MCP server, which provides ranked failure lists, change timelines, and filtered logs. The study measured diagnosis, not remediation, and each arm ran each fault only once. However, the per-fault data is public, allowing teams to replay incidents. The recommendation is clear: install Radar (open source, Apache-2.0) or a similar context provider, point your existing agent at a read-only MCP endpoint, and measure the baseline before committing to new vendors.