Stop burning cash on GPT-4 for tasks that barely require intelligence. If your agent architecture treats every step like a frontier-model problem, you are bleeding capital. New insights from Intercom and Superhuman Mail prove that narrowing high-volume steps to smaller, post-trained models can save hundreds of thousands of dollars a month without degrading user experience. The trick isn't just picking a cheaper model; it's identifying which specific jobs in your agent pipeline are narrow enough to survive the downgrade.

The Intercom Case Study: Canonicalization

Fergal Reid, Chief AI Officer at Intercom, revealed on the Chain of Thought podcast that their query canonicalization step was costing them approximately $250,000 a month on GPT-4.1. This step, which rewrites colloquial user queries into canonical forms for better retrieval, was running on a model far more powerful than necessary. By replacing GPT-4.1 with a post-trained 14-billion-parameter Qwen model, Intercom saved almost all of that several-hundred-thousand-dollar monthly inference bill. The key was proving the step worked on the strongest model first, then moving that single, high-volume job to a specialized, cheaper alternative.

Superhuman's Approach to Auto-Labeling

LoΓ―c Houssier, CTO of Superhuman Mail, follows a similar philosophy but starts with the best model to ensure feature viability. For their auto-labeling system, which classifies incoming emails into categories like 'FYI', 'pitch', or 'newsletter', they initially used a strong model on every email. Once the team identified the 10 to 12 useful labels, the task transformed from a complex reasoning problem into a typical BERT-classifier job. They fine-tuned an open model for this specific classification task and moved it to cheap, dedicated infrastructure, keeping the heavy agentic reasoning on the latest Opus model.

How to Validate the Swap

You cannot just swap models and hope for the best. Intercom validates changes using two specific metrics: soft resolutions (user leaves without replying) and hard resolutions (user confirms the issue is resolved). Reid notes that hard resolutions typically occur at 30 to 40 percent of the soft-resolution rate. When swapping models, the goal is to keep hard resolutions steady while potentially increasing soft resolutions. If the ratio shifts, the product is likely deflecting users rather than solving their problems. For Fin, Intercom's customer service agent, this validation process involves A/B tests on millions of interactions to catch even a tenth-of-a-percentage-point deviation.

Key Takeaways

  • Identify narrow, high-volume steps like canonicalization or fixed-label classification as candidates for cheaper models.
  • Always prove the feature works on the frontier model first to establish a quality baseline.
  • Rank agent steps by monthly spend, including all retries, to find the biggest cost drivers.
  • Use post-trained smaller models, like the 14B Qwen, for specific tasks where general intelligence is overkill.
  • Monitor hard resolution rates to ensure cost-cutting does not degrade actual user satisfaction.

The Bottom Line

Stop paying frontier prices for classification problems. If a step is narrow, high-volume, and has a scoreable output, it belongs on a cheap, specialized modelβ€”not Opus.