The developer known as kyalo460 has released AiWatchdog, an open-source monitoring tool designed to solve the 'black box' problem of autonomous coding agents. Hosted at aiwatchdog.vercel.app and licensed under MIT, the tool provides real-time visibility into agents like Kilo, Claude Code, Codex, and Aider. The core value proposition is distinguishing between an agent that is legitimately busy compiling or testing and one that is wedged in an infinite retry loop burning tokens while the developer is away.

The Detection Engine: Weighted Signal Ensembles

The naive approach to stuck detectionβ€”alerting on five minutes of silenceβ€”fails because agents naturally go quiet during dependency installation or Docker runs. AiWatchdog instead employs a weighted ensemble of independent detectors, each producing an explainable signal. For instance, 'No heartbeat' carries a weight of 20, while 'Repeated similar errors' carries a heavier weight of 25. The system calculates confidence as the clamped sum of these triggered weights, ensuring that total silence scores 55β€”five points below the 60-point threshold required to trigger a 'STUCK' alert. This design rule ensures that silence alone is never considered evidence of failure.

Architecture and Implementation Details

The tool integrates via a simple command wrapper, such as watchdog wrap -- kilo run "fix the failing test", which passes stdin through and propagates exit codes while sending heartbeats every 15 seconds. The backend is built with TypeScript and Python SDKs, featuring a pure detection engine that avoids database or network dependencies to maximize testability. The project includes approximately 1,500 tests with no mocks in end-to-end scenarios, utilizing real API processes, Postgres databases, and WebSocket connections. Notably, the developer opted against using Next.js or an ORM, favoring hand-written SQL to maintain deterministic control over monthly RANGE partitioning and compare-and-set state writes.

Known Limitations and Roadmap

The release acknowledges significant current limitations, particularly regarding editor-based agents. The VS Code extension cannot see chat traffic inside other extensions like Cline or Roo, meaning the localhost bridge and terminal signals represent the current ceiling for observability in those environments. Furthermore, the capability flags supportsPauseResume and supportsProcessMetrics are currently false for all adapters. The roadmap includes extracting the sweeper to a BullMQ worker with advisory locks to prevent double-sweeping across API replicas, along with adding SSO, MFA, and audit logs.

Key Takeaways

  • AiWatchdog uses a weighted signal ensemble to avoid false positives from legitimate long-running tasks.
  • The tool is MIT-licensed and supports major agents including Kilo, Claude Code, Codex, and Aider.
  • Detection logic is pure and exhaustively tested with ~1,500 tests using real infrastructure.
  • Current limitations include lack of visibility into internal chat traffic of VS Code extensions.

The Bottom Line

This is a pragmatic solution to a universal pain point in agentic workflows, prioritizing honest capability matrices over feature bloat. The open-source nature of the detection methodology invites community refinement, which is critical for a tool that must adapt to evolving agent behaviors.