If you are running a tool-using AI agent, your monthly bill is likely being eaten alive by a single, non-obvious mechanic: the quadratic growth of conversation history. A recent deep dive into token mathematics reveals that a mere four-message support chat can consume over 45,000 input tokens. This is roughly twelve times the 3,700-token estimate Anthropic provides for standard support conversations. The culprit isn't a bug; it's the architecture. Each turn forces the model to re-read the entire system prompt, tool definitions, and every previous message in the chain.

The Quadratic Reality of Context

The core issue is that history cost scales with the square of the number of turns. The formula provided demonstrates that moving from four turns to eight multiplies history tokens by 4.7, not 2. In the worked example, a small-business support bot handling 3,000 conversations a month sees 44% of its model bill driven purely by resending earlier turns. This means that optimizing your system prompt is secondary to managing conversation length. If your agents are looping or if users are engaging in long back-and-forths, you are paying exponentially more for marginal utility.

Pricing Across the Landscape

When priced against six current models as of October 2026, the same workload ranges from $17 to $342 a month in model fees alone, before caching. OpenAI's gpt-6-luna sits at the low end of $17.09, while Claude Sonnet 5.5 and OpenAI gpt-6.1-sol hit the high end of $341.88. However, prompt caching is the biggest lever you can pull. Enabling caching on Sonnet 5.5 cut the bill by 59%, dropping it to $140.79. But beware: caches expire in five minutes by default, and a quiet customer breaks the chain, forcing an expensive cache write on the next interaction.

Hidden Infrastructure Costs

Token prices are only part of the equation. The report highlights that vector databases, monitoring, and search fees add significant overhead. Pinecone's Standard plan starts at $50 a month, while Langfuse monitoring costs $29 for 100k units. Furthermore, self-hosting open models often looks cheaper on paper but fails at small-business volumes. Renting a single NVIDIA A10 GPU on Lambda costs about $942 a month, which is more than the highest API bill in the example. Self-hosting only makes sense if you have high volume, strict data privacy requirements, or need a specialized fine-tuned model.

Key Takeaways

  • History resending accounts for 44% of model costs in typical support agents.
  • Input-to-output ratios are skewed heavily toward input (38:1 in the example), making input pricing critical.
  • Prompt caching can reduce bills by 59%, but requires careful management of cache expiration and minimum token thresholds.
  • Self-hosting is rarely cost-effective for small businesses, often costing more than the most expensive API tier.

The Bottom Line

Stop obsessing over model selection and start ruthlessly managing conversation length. The architecture of agentic loops means that history is the enemy of your budget, and caching is your only shield against quadratic inflation.