The hype cycle around Generative Engine Optimization (GEO) just hit a brick wall. A new experiment conducted by Joe Shirey, an engineer at Google Cloud, demonstrates that traditional SEO tactics have minimal influence on which products AI coding agents actually recommend. Despite the industry’s frantic pivot to optimizing content for LLMs, the data suggests that an agent’s choices are overwhelmingly dictated by its underlying training weights, not the search results it scrapes in real-time.

The Death of the Search-First Assumption

Shirey’s research was sparked by a shift in developer behavior, mirroring consumer trends where 46% of AI users now start product research on platforms like ChatGPT or Claude rather than search engines. To test if vendors could game these systems, he ran 5,580 trials using Claude Sonnet 5 to answer 310 architecture questions. He employed a custom MCP server to intercept and bias search queries, forcing Google Cloud products to the top of the results. The result? Even with search results heavily rigged in favor of GCP, the primary selection rate only climbed from 9.2% to 17.1%.

Training Data Is the Real Kingmaker

The most critical finding is that AI agents rarely use search tools unless explicitly forced to. In the baseline test, Claude Sonnet 5 searched the web in only 0.3% of trials, answering almost entirely from its training data. When search was available but not encouraged, biased search results had virtually no impact, changing selection rates by just 1 percentage point. This indicates that what a model learned during its training phase—shaped by years of documentation, tutorials, and community adoption—is the dominant factor in its recommendations, not the content of a live search query.

Key Takeaways

  • Search Frequency is Low: By default, agents like Claude Sonnet 5 rarely search the web for advisory questions, relying instead on internal knowledge.
  • SEO Has a Ceiling: Even under ideal conditions with forced searching and biased results, SEO only increased product selections by ~8 percentage points.
  • Mentions ≠ Selections: Rigged search results significantly increased product mentions (from 56% to 89%) but had a much weaker effect on actual recommendations.
  • Vendor Control is Illusory: Developers control the prompt and tool configuration, meaning vendors cannot force an agent to use their preferred search biases.

The Bottom Line

Stop wasting money on short-term GEO hacks; your product’s reputation in the training corpus is what actually matters, so focus on long-term community trust and high-quality documentation.