Despite the hype around autonomous coding agents, a recent deep-dive by EmmFlow reveals that AI still lacks the 'taste' and contextual awareness required for complex user experience design. The author, working on Emm—a local-first notes manager and issue tracker—attempted to offload the design of a multi-step workflow involving more than 20 screens to AI agents. The result? A cautionary tale of wasted compute cycles and frustrated developers that underscores a critical truth: current LLMs cannot replace human-led prototyping for intricate back-end machinery.
The Trap of Full-Stack Autonomy
The first strategy involved trusting an agent named Astra to handle the entire lifecycle: asking questions, identifying decisions, generating specs, and implementing code. While Astra successfully handled the plumbing that would have previously taken days, it failed fundamentally at UX optimization. The agent created a flow requiring seven clicks for actions that only needed four, and it navigated users to pages where no decisions were being made. This highlights a persistent gap in current AI agents: they can write code, but they cannot inherently judge efficiency or user intent.
Why Iterative Screen-Building Fails
Attempting to regain control, the developer switched to building the experience one screen at a time, explicitly dictating each step to the AI. This approach collapsed under the weight of technical constraints and the iterative nature of design. Because Emm is built in pure Rust without a heavy framework, simple changes took agents 10-20 minutes to compile and implement. Adding a single branch to a login flow took hours. Furthermore, design is inherently iterative; the initial 'working' build was confusing and edge-case poor, requiring another 8+ hours to rebuild using the same slow method.
The HTML Prototype Breakthrough
Success finally arrived when the developer abandoned code-first generation in favor of visual prototyping. By switching to ChatGPT Desktop, the developer asked the agent to build a mock application in HTML that spanned different surfaces—including web browsers, the native app, and macOS passkey pop-ups. This approach allowed for rapid iteration and testing of different starting configurations via a dropdown menu. Although the process still required three hours of human decision-making and testing, it produced a crystal-clear prototype that covered all use cases before a single line of production code was written.
Key Takeaways
- Context is King: AI agents struggle to maintain context across 20+ screen workflows, leading to inefficient user paths and redundant navigation.
- Tech Stack Matters: In pure Rust or low-framework environments, AI implementation cycles are too slow for iterative design; prototyping in HTML bypasses compilation bottlenecks.
- Separate Design from Implementation: Use AI to generate visual prototypes (HTML) for rapid feedback, then implement the validated design in code in one go.
The Bottom Line
Stop letting AI agents drive your UX decisions. Use them for boilerplate and rapid HTML mockups, but keep the human in the loop for workflow logic until the spec is perfect.