In the rush to slap "agent" labels on every LLM wrapper, we are often paying a premium for autonomy we don't need. A new experiment from calm.rocks puts hard numbers on this trade-off, revealing that for predefined workflows, letting the model decide the next step costs 5.6 times more tokens while delivering virtually the same accuracy as a simple code pipeline.
The Experiment: Pipeline vs. Agent
The study compared two approaches to extracting four specific fields (parties, date, value, termination notice) from 20 synthetic vendor contracts. Both methods used the same model, openai/gpt-oss-120b on Groq, with temperature set to 0 to ensure determinism. The "Pipeline" approach used a single, fixed-schema prompt. The "Agent" approach used a tool loop capped at 8 turns, equipped with a find_line search tool and a calculate arithmetic evaluator. The hypothesis was that the agent's dynamic control flow would justify its complexity in handling messy data.
Accuracy Tied, Costs Skyrocket
The results were stark. Across three runs, the pipeline missed 1 field per run, while the agent missed 2. Both achieved approximately 98-99% field accuracy. However, the resource consumption was wildly different. The pipeline averaged 356 tokens per contract, while the agent averaged 1,981 tokensβa 5.6x increase. In the "math" group, where annual values had to be computed, the agent spent 3,039 tokens per contract compared to the pipeline's 346, an 8.8x gap. The calculator tool removed the arithmetic burden but added the overhead of reasoning, tool call emission, and re-sending growing conversation histories.
The Hidden Cost: Unstable Resource Consumption
Beyond the average cost, the agent introduced significant variance. While answers remained stable due to temperature=0, the token count per contract fluctuated dramatically. The pipeline's largest run-to-run difference was 7 tokens. The agent's was 1,848 tokensβnearly the size of its average cost. This instability stems from dynamic control flow: if the model takes an extra turn, it re-sends the entire conversation history. For production systems, this means budgeting, latency forecasting, and rate limit management become nightmare scenarios when the "next step" is decided at runtime rather than design time.
When Is Autonomy Actually Worth It?
The author argues that control flow should live in code when the sequence of operations is known at design time. Agents are only compelling when the next step depends on information discovered at runtime, such as in open-ended research or iterative debugging. For the contract extraction task, the path was linear and predictable. The agent essentially paid a 5.6x tax to rediscover a known path on every single document. The lesson is clear: add agentic complexity only when it demonstrably solves a problem that fixed workflows cannot.
Key Takeaways
- Fixed Tasks = Fixed Pipelines: If you know the steps, hard-code them. Don't let the model decide.
- Token Overhead is Structural: Agent loops re-send growing history, causing exponential cost increases with each turn.
- Cost Variance is a Risk: Dynamic control flow leads to unpredictable token usage, complicating production budgeting.
- Tools Aren't Free: A calculator tool helps weak models but adds significant overhead for capable ones.
The Bottom Line
Stop wrapping every simple task in an agent loop. If your workflow is linear, your code should be too. Autonomy is a feature, not a default.