Sulcus has opened early access to its platform for observing and controlling AI agents, addressing a critical pain point for developers: the lack of visibility once an agent is live. The tool supports major frameworks including LangGraph, CrewAI, and the OpenAI Agents SDK, allowing users to monitor token usage, model calls, and tool activity. Unlike standard tracing tools that only record history, Sulcus provides active control mechanisms, enabling users to stop runs or approve specific actions in real time.

Three Tiers of Integration

The platform offers three distinct integration modes depending on where the agent runs. For hosted Python runs, Sulcus executes the agent from a public Git repository, providing full control to start, stop, and rerun processes with configurable CPU, memory, and time limits. For agents running on your local machine, such as Claude Code and Codex, Sulcus acts as a middleware via a command-line tool. This allows users to approve or deny permission requests for file changes and commands, though ending observation does not stop the local agent itself. For agents deployed in other environments, like Google ADK or Microsoft Copilot Studio, Sulcus operates in a read-only mode. It ingests telemetry data, such as OpenTelemetry spans from ADK or Application Insights from Copilot Studio, to display run history after the fact. In this mode, Sulcus cannot intervene in the execution flow; it purely visualizes agents, sub-agents, and tool calls to help debug issues without altering the live service.

Privacy and Token Limits

Sulcus emphasizes privacy by not storing prompt text, model responses, or tool inputs and outputs for its framework integrations. However, for hosted runs, any output printed to stdout or stderr is recorded in the run history, meaning developers must avoid printing sensitive data. The platform includes a token limit guard that blocks subsequent model calls once a reported threshold is reached, though this is not a hard billing cap since in-flight requests may still complete and incur costs from the provider.

Key Takeaways

  • Sulcus offers three integration tiers: full control for hosted runs, approval-only for local agents like Claude Code, and read-only telemetry for external services.
  • The platform prioritizes privacy by excluding prompt text and model outputs from storage, except for stdout/stderr logs in hosted environments.
  • Token limits act as behavioral guards rather than hard billing caps, potentially allowing in-flight requests to exceed thresholds.

The Bottom Line

Sulcus fills a necessary gap in the AI agent stack by moving beyond passive tracing to active intervention, though its utility is strictly tiered by where your code executes.