A developer going by the handle 'agentic_architect' on DEV.to recently ran what they called a little experiment with Claude Opus 5, and the results should serve as a reality check for anyone banking on AI agents to handle complex, open-ended research tasks without close supervision. The task was straightforward enough: scan the web landscape, dig into niche corners of the internet, and identify viable low-competition opportunities for making money online. Sounds simple in theory, right? Wrong. Within hours, Opus 5 had torched through the author's entire £20 monthly Cursor budget and produced 44 files of what they described as 'absolute garbage.'
What Went Wrong
The experiment highlights a fundamental issue with current AI agent capabilities when faced with genuinely open-ended tasks that lack clear constraints. Opus 5 was apparently given free rein to explore, research, and synthesize findings across the web, but without guardrails or specific success criteria baked in from the start. The result wasn't actionable intelligence about money-making opportunities—it was a mess of incoherent outputs that added up to nothing useful. This isn't necessarily a knock on Claude Opus 5 specifically; it's more an indictment of how we approach agentic AI for complex workflows. When you give a model an underspecified task and let it run unsupervised, you're essentially asking for trouble.
The Budget Problem
What's particularly striking here is the economic angle. £20 might not sound like much, but when Opus 5 burns through that budget in a single experiment producing nothing of value, you've got to ask yourself: what are we actually paying for? Cursor's integration with Claude models has been marketed as a productivity multiplier for developers, yet this experiment suggests that uncontrolled agentic behavior can turn expensive fast. The author didn't set spending limits or checkpoints—they just let Opus 5 run until the well was dry. That's on them, sure, but it also exposes how easy it is to lose control of AI spend when you're not actively monitoring outputs against costs.
Key Takeaways
- Define clear success criteria and constraints before launching any agentic task—no open-ended research without guardrails
- Set budget alerts and automatic stop conditions for long-running tasks to avoid surprises on your invoice
- Open-ended exploration tasks require human-in-the-loop checkpoints, not autonomous execution from start to finish
- The gap between 'impressive demo' and 'reliable production tool' remains significant for AI agents handling complex workflows
The Bottom Line
This experiment is a useful reminder that AI agents aren't magic. Opus 5 might be the most capable model on paper, but without proper task scoping and oversight, you'll just burn money and get gibberish. For developers looking to actually ship value with these tools, start small, define success metrics upfront, and never let an agent run unsupervised until you've validated it can handle simpler tasks first.