The promise of AI-accelerated development often collapses under the weight of vague requirements and architectural drift. However, a new logbook from senior engineer sidiar offers a pragmatic blueprint for maintaining speed without sacrificing cohesion. By building a 45,000-line React and TypeScript application, "Game of Life Studio," using the BMad method, the developer demonstrated that Spec-Driven Development (SDD) can indeed bridge the gap between executive demands for three-week delivery and the reality of production-grade code. The project, live at game-of-life-studio.com, serves as a stress test for AI agents, revealing that while code generation has become nearly instantaneous, the human role has shifted from writing logic to orchestrating and validating it.

The Planning Phase: Precision Comes at a Cost

The experiment began with a rigid adherence to the BMad framework, which mandates a Project Brief, Product Requirements Document (PRD), and Architecture spec before any implementation begins. This approach forced the developer to resolve every edge case and contradiction upfront, a process that proved both a strength and a friction point. While AI agents excelled at detecting missing requirements, the mandatory resolution of even low-priority issues created a bottleneck, preventing the natural deferral of minor decisions that often characterizes agile development. Furthermore, the communication style of the AI became a significant hurdle; the agents produced verbose outputs filled with internal references like "AR-33" and poetic terminology, requiring the developer to spend weeks simply learning to parse the machine’s "employee" persona rather than focusing on the code itself.

Execution: Automating the Workflow, Not the Decision

The turning point in the project came when the developer stopped treating AI as a passive coder and started building custom orchestration skills. Initially, the execution phase was sluggish, with the developer bogged down by manual reviews of every pull request. The solution was the creation of an "implement next story" skill, which automated the pipeline: spawning subagents to create a story, implement it, and review it on a different model, all within fresh sessions. This separation of concerns was critical. The developer discovered that using the same model for both implementation and review resulted in silent failures where the AI effectively graded its own homework. By enforcing a cross-model review policy—such as having Sonnet implementations reviewed by Opus—the system caught errors that self-review missed, turning the AI agents into reliable, though verbose, junior developers.

What Broke: Parallelization and Silent Failures

As the project scaled, parallelization introduced new classes of failure. The developer learned that running two AI agents in the same directory led to race conditions where commits from one agent contaminated the work of another. This necessitated the adoption of git worktrees to isolate lanes, a fundamental infrastructure shift for AI-driven workflows. Another critical failure involved the metrics themselves; initial stats showed a 65-hour review time, which was actually a measurement error counting idle time and usage-limit resets. This highlighted a deeper truth: the slowest part of the pipeline was no longer the AI’s generation time, which averaged 67 minutes per story, but the human latency in approving PRs. The AI could write code in minutes, but the developer’s decision-making cycle remained the primary constraint on delivery velocity.

Key Takeaways

  • AI agents are highly efficient at detecting missing requirements and edge cases in specifications, but they produce verbose, hard-to-read outputs that require new communication skills from humans.
  • Cross-model review is essential; using the same model for implementation and review creates a feedback loop that misses errors, whereas distinct models (e.g., Sonnet implemented, Opus reviewed) catch issues effectively.
  • Parallelization requires strict isolation; running multiple AI agents in the same directory leads to state contamination, necessitating the use of git worktrees or similar isolation techniques.
  • The bottleneck in AI-assisted development has shifted from code generation to human review; the AI can produce a PR in under an hour, but the time to human decision-making is the new rate limiter.

The Bottom Line

Stop asking how fast AI can write code and start asking how fast you can decide if the code is right. The tools are ready; your review process is the bottleneck.

Infrastructure Implications

For teams looking to adopt SDD, the logbook suggests that infrastructure must evolve to support AI orchestration. The developer’s custom skills and git worktree setup are not just conveniences but necessities for maintaining velocity. The 82.6 million cache-read tokens used across thirty stories demonstrate that cost efficiency is achievable through smart architecture, but only if the human-in-the-loop process is streamlined. The era of "vibe coding" without structure is over; the future belongs to those who can specify, orchestrate, and review with precision.