Building a production-grade React Native app with AI agents is less about code generation and more about architectural governance. A developer building ParkEase, a peer-to-peer parking marketplace for India, detailed how using Claude Code resulted in 72 commits and 22,800 added lines, but crucially, the discovery of six hidden bugs in already-merged code. The project leveraged a custom skill called /flow, a routing layer that enforces strict gates for thinking, planning, building, and reviewing, ensuring that AI outputs are treated as hypotheses subject to rigorous inspection rather than trusted artifacts.
The /flow Architecture and Decision Ledger
The core of this workflow is the /flow skill, which acts as a governance layer rather than a simple prompt. It routes tasks through phases like 'Think' and 'Plan,' with mandatory gates such as 'no code until design approval' and 'test-driven development.' A key rule states that 'the governance chain is read from the docs, never inferred,' meaning Architecture Decision Records (ADRs) override all other instructions. This hierarchy proved vital early on when the AI rejected a suggested UI color scheme because it conflicted with a previously established ADR, demonstrating that the system can enforce project consistency over its own generative tendencies.
Uncovering Critical Defects in Merged Code
The most significant outcome was not the new features, but the detection of six critical defects in code that had already passed standard reviews. These included a broken proof-photo upload for valets due to a mismatch between multipart/form-data and JSON expectations, a driver wash screen that failed for all real partners because of a null schema issue, and a redirect error that left valet and washer partners with an 'Unmatched Route' upon opening the app. The AI's multi-lens review stack, running security, database, and React specialists in parallel, identified these issues by treating the entire branch as a subject for adversarial testing.
Opus 5.5 and the Value of Judgment
The developer highlighted that Opus 5.5 excelled not in raw coding speed, but in 'judgment in the controller seat.' The model frequently admitted when its own plan contained bugs, such as an idempotency key error that would have caused upload failures, and verified assumptions before acting on reviewer feedback. This willingness to self-correct and push back on scopeβrefusing to implement a sparkline feature that would have required risky client-side money arithmeticβshows a maturation in agentic workflows where the AI acts as a senior engineer who documents their reasoning.
Token Optimization and Rate Limit Reality
Despite the technical successes, the process was slowed significantly by rate limits, with the Sonnet weekly quota exhausted on day one. The developer implemented a token-optimization playbook that included handing over artifacts as files rather than pastes, using short return contracts to minimize context bloat, and tiering models with Haiku for transcription and Opus for judgment. This approach, combined with a local knowledge graph for codebase queries, allowed the team to resume interrupted agents without losing context, though the total wall-clock time still stretched across two days.
Key Takeaways
- AI agents are most effective as a 'team' of specialized reviewers rather than a single code generator.
- Strict governance layers like /flow prevent AI hallucinations from overriding established architectural decisions.
- Multi-lens reviews can uncover critical bugs in already-merged code that single-pass tests miss.
- Token optimization strategies like file-based artifact handoff are essential for long agentic sessions.
- The primary value of AI development is inspectability and decision documentation, not necessarily speed.
The Bottom Line
The future of AI-assisted development isn't about faster typing; it's about creating systems where every decision is traceable, every bug is surfaced by adversarial review, and the human developer retains final authority over architectural truth.