The consensus among senior engineers is shifting: the problem with AI coding agents isn't that they break the build. It's that they quietly rot the architecture over time. A new deep-dive from developer jgauffin on DEV.to argues that while individual agent commits look cleanβ€”tests green, diffs reviewedβ€”the cumulative effect is a codebase where the implementation becomes the only record of intent. This 'silent decay' creates technical debt that is invisible until a refactor goes catastrophically wrong. The proposed solution? A VS Code extension called Kiwipow Agent, which attempts to decouple product intent from code execution.

Decoupling Intent From Implementation

The core philosophy of Kiwipow Agent is that code should never be the source of truth for what a product does. The tool's feature planner operates entirely outside the codebase, generating named rules in plain language based on documentation and historical specs. This creates an artifactβ€”a 'feature file'β€”that a product manager can read, argue with, and contradict. By keeping the spec separate from the code, the system ensures that future changes are planned against what the product *should* do, not just what it currently does. This prevents the common failure mode where agents optimize for existing, potentially flawed, behavior.

Tests That Defend Promises, Not Snapshots

Perhaps the most critical innovation here is the shift in how tests are generated. Traditional agent-written tests are 'snapshots' of the current implementation; they break the moment you refactor because they assert *how* the code works, not *why* it works. Kiwipow Agent generates tests directly from the approved rules. If a rule states 'Refund on cancel,' the test fails when the refund logic breaks, regardless of how the underlying services are restructured. This approach claims to restore the actual purpose of a test suite: providing the safety net needed to restructure code without fear. It turns a red test into a named broken promise, rather than a cryptic line number error.

Bounded Autonomy and Cognitive Limits

The tool also addresses the runaway cost of 'continuous improvement.' Instead of allowing an agent to rewrite code indefinitely, Kiwipow Agent applies strict, user-defined limits on cognitive complexity and file size. When a feature's tests pass, the implementation files are measured against these metrics. If a file exceeds the limit, the agent proposes a split, but the human decides whether to execute it. Crucially, the agent is allowed only one pass per feature to avoid 'moving code sideways'β€”a common issue where models churn through working code without adding value. This restraint is designed to keep the cost of quality improvement predictable and bounded.

Key Takeaways

  • Silent Decay is Real: Agent-written code often passes immediate checks but accumulates unenforceable rules and bloated files that hinder future development.
  • Specs as Source of Truth: Kiwipow Agent enforces a workflow where plain-language specs drive planning, preventing agents from optimizing for legacy bugs.
  • Behavioral Testing: By generating tests from rules rather than implementation, the tool aims to make refactoring safe again.
  • Human-in-the-Loop Limits: The system uses strict, bounded passes for code quality improvements to prevent infinite, costly churn.

The Bottom Line

If your agent is just writing code, you're already paying for this decay. Kiwipow Agent is a necessary intervention for teams tired of debugging their own AI's architectural shortcuts. It’s a reminder that 'working code' is a low bar for agents that are supposed to be engineering partners, not just autocomplete on steroids.