If you are still budgeting AI agents by multiplying token counts by price-per-million, you are optimizing for the wrong variable. In a new post on DEV.to, developer Cleo Cliona argues that per-token pricing is a unit cost, not a task cost, and it is fundamentally misleading for anyone trying to determine if their agent stack is actually cheap to run. The core insight? A model that is expensive per token but nails a task in one turn is vastly cheaper than a budget model that hallucinates its way through ten failed attempts.
The Hidden Tax of Failure
Cliona describes a personal reckoning after realizing that token metrics told her almost nothing about actual value delivery. She began tracking 'cost per completed task' instead, logging turns, tool calls, and success status after every session. The data revealed a brutal truth: failed runs consume the exact same budget as successful ones. When an agent enters a loop—re-reading a file, making the same bad edit, hitting the same error—it burns through credits without producing a usable result. Those tokens aren't free just because the output was garbage; they are the most expensive tokens in the system because they bought nothing.
Routing Beats Raw Power
The analysis also dismantles the obsession with choosing the 'best' model. Cliona found that routing strategy—how tasks are broken down, constrained, and distributed—impacts cost more than the underlying model choice. A well-routed task sent to a mid-tier model often finishes faster and cheaper than a poorly-routed task sent to a top-tier model. When instructions are vague, even the smartest agent wanders, accumulating dead-end turns. When tasks are explicitly constrained and sub-divided, smaller models execute efficiently, reducing the total turn count and, consequently, the total bill.
Key Takeaways
- Per-token pricing is a vanity metric; cost-per-completed-task is the only budgeting metric that matters.
- Failed agent runs incur the same input costs as successful ones, making debugging loops financially catastrophic.
- Task routing and explicit constraints reduce total turn counts more effectively than upgrading to larger models.
- Logging success/failure status is more valuable for cost analysis than logging raw token consumption.
The Bottom Line
Stop looking at your token usage dashboard like it’s a bank statement. If you aren't tracking success rates per task, you are just measuring how fast you burn money on hallucinations. Fix your prompts, constrain your loops, and stop pretending a cheaper model is cheaper if it takes four times as many turns to do the same job.