The era of micromanaging AI coding agents is ending for some power users. Alexey Indeev, an engineer at Spare, recently shared a radical shift in his development workflow: he stopped reviewing his agents' code line-by-line. In just two weeks, this change allowed him to merge 562 pull requests while still finding time for four nights of hiking. The key wasn't better prompts alone, but a robust system of automated guardrails and a philosophical shift from reviewing code to reviewing intent.

The Guardrail-First Architecture

Indeev's approach relies on the premise that AI code is non-deterministic, but so is human code. He argues that we already build SRE systems to prevent human errors; now, those systems need to be beefed up for agents. Every PR in his internal platform, Sightline, must pass unit tests, end-to-end tests, CI checks, and automated reviews from Bugbot and Strix before entering the merge queue. Crucially, the guardrails are self-healing: if a bug slips through, the team fixes the guardrail, not just the code, creating a feedback loop that hardens the system over time.

From Tickets to Outcomes

The workflow centers on a 'coordinator' agent that manages sub-agents, keeping the top-level context clean. Instead of assigning specific tickets like 'add a role editor,' Indeev sets high-level goals such as 'make Sightline's permissions work like Spare's everywhere.' He explicitly prompts the coordinator: 'Accomplish the plan. Ship code using auto-merge. You're the coordinator. Don't do work directly.' This allows the agent to parallelize work, feature-flag changes, and ship in small, verifiable increments. One notable example involved a goal running for over a day, where the agent iterated until it hit the desired outcome, with Indeev only checking in via HTML progress reports.

Token Bottlenecks and Multi-Account Swapping

This level of autonomy burns through tokens at an alarming rate. Indeev noted that individual plans no longer provide enough credits, leading to expensive overages. To mitigate this, he runs multiple Claude accountsβ€”about five at roughly $200 eachβ€”and uses a tool called 'Claude Swap' to automatically switch accounts when usage hits 90%. While effective, he admits this method is clunky, causing loss of artifacts and requiring constant re-logins. He is currently evaluating alternatives like Paseo, Orca, and Superset to streamline the multi-account management process without the friction.

Key Takeaways

  • Review plans and intent, not code: Spend time aligning with the agent on the 'what' and 'why' before letting it execute.
  • Guardrails are the job: Automated tests, CI, and review bots must be rigorous enough to catch agent errors automatically.
  • Ship small and often: Auto-merge with feature flags allows for safer, incremental deployments compared to large, local iterations.
  • Set outcomes, not tickets: Give agents business problems (e.g., 'fix slow CI') rather than specific coding tasks to leverage their reasoning capabilities.
  • Manage token costs proactively: Multi-account swapping tools are necessary for high-volume autonomous workflows.

The Bottom Line

Indeev’s experiment proves that the bottleneck in AI-assisted development is no longer the agent's coding ability, but the human's ability to define clear intent and build robust safety nets. We are moving from being coders to being architects of self-healing systems.