A developer going by rubenflamshepherd has launched a blind taste test platform that pits Claude against Codex in a head-to-head design generation showdown—and the results might surprise you. The experiment, posted to Hacker News on July 25th, asks visitors to guess which AI model generated each book landing page without knowing the answer until after they vote. The project emerged from what began as a simple experiment: asking Codex to create a companion website for a book the developer was enjoying. To their shock, the output was "actually good." That unexpected success sparked a deeper investigation into how different AI models approach visual design and layout generation when given identical prompts. "I went down a bit of a rabbit hole prompting different models to generate landing pages for different books," rubenflamshepherd explained in their Show HN post. The resulting platform lets visitors compare side-by-side designs without knowing which came from Anthropic's Claude or OpenAI's Codex until they've made their selection.

How the Blind Test Works

Visitors to taste.rubenflamshepherd.com are presented with two generated landing pages for various books. Users must choose which design they prefer before discovering which AI model created each option. The setup removes bias by eliminating brand recognition—visitors judge purely on aesthetic merit and functionality rather than favoring a particular company's reputation. The platform currently features multiple book titles, each with designs generated by both Claude and Codex using consistent prompting strategies. This methodology ensures that differences in output stem from the models' inherent design philosophies rather than variations in how instructions were given to each system.

What This Reveals About AI Design Agents

Beyond the entertainment value of guessing which model made what, the experiment offers genuine insight into the current state of AI-driven design generation. Both Claude and Codex have positioned themselves as capable coding and creative assistants, but their approaches to visual layout and aesthetic choices appear distinctly different. The taste test format also highlights a practical question for developers: when should you reach for one model versus another? If outputs are genuinely indistinguishable to casual observers, the choice might come down to factors like API pricing, integration complexity, or specific domain strengths rather than pure design quality.

Key Takeaways

  • Codex surprised its creator with "actually good" output on a real project before formal testing began
  • The blind format removes brand bias and focuses purely on visual merit
  • Design differences between models exist but may not be obvious to end users
  • Multiple book landing pages are available for comparison on the platform

The Bottom Line

This isn't just a parlor trick—it's a glimpse at how far AI design generation has come. When humans can't reliably distinguish outputs from competing models, we've crossed a threshold. The real question isn't which AI is better at design anymore; it's whether we even need human designers for commodity landing pages in 2026.