Yotam Yemini, founder of Causely, has laid out a definitive operational framework for AI agents in production, arguing that shipping an agent is only half the battle. The core thesis, published on October 5, 2026, posits that reliable agent operations require a distinct three-layer stack: instrumentation, evaluation/observability, and causal reasoning. This mirrors the evolution of microservices observability a decade ago, but with the added complexity of non-deterministic LLM behavior.

Layer One: The Instrumentation Foundation

Everything begins with capturing the agent's actual behavior, not just its intended behavior. Yemini emphasizes adopting OpenTelemetry's GenAI semantic conventions, which provide a standardized vocabulary for LLM calls, tool usage, and token consumption. Although these conventions are still under active development as of mid-2026, requiring teams to budget for attribute renames, they have consolidated enough to serve as the primary data source. This layer creates the structured, correlated record necessary for any subsequent analysis.

Layer Two: Evaluation and Observability Platforms

Once signals are captured, the second layer determines if the agent is performing correctly. This domain is dominated by platforms like Comet, Langfuse, Braintrust, Galileo (now under Cisco), and Arize (acquired by Dynatrace for a reported $915M). These tools focus on scoring outputs, catching regressions, and monitoring operational health metrics such as latency, error rates, and redundant tool calls. Yemini notes that while these platforms answer 'what' the agent did and 'whether' it was good, they often fail to explain 'why' a degradation occurred.

Layer Three: Causal Reasoning Across the Stack

The third and most critical layer addresses root cause analysis by connecting agent behavior to the underlying infrastructure. Agents rely on a complex chain of services, vector stores, APIs, and databases. When performance slips, the cause is often external to the agent itself, such as a stale embedding index or a database under memory pressure. Yemini argues that correlation-based approaches, while improved by recent acquisitions consolidating telemetry data, hit a ceiling because they cannot distinguish signal from noise in constantly changing environments.

The Causely Approach: Deterministic Over Probabilistic

Causely differentiates itself by maintaining a live causal model of the environment rather than relying solely on machine learning correlations. This deterministic approach aims to eliminate hallucinated root causes, providing specific, actionable diagnoses. Yemini cites internal customer metrics claiming 63% faster incident resolution and 57% lower operating costs. A notable example involves Causely proactively detecting degradation and triggering background agents to submit twelve pull requests, effectively acting as an autonomous engineer dedicated to system health.

Key Takeaways

  • Instrumentation via OpenTelemetry GenAI conventions is the non-negotiable first step for any production agent.
  • Evaluation platforms like Langfuse and Arize are essential for quality scoring but insufficient for root cause analysis.
  • The $915M Dynatrace acquisition of Arize signals industry consolidation around unified agent and infrastructure telemetry.
  • Causal reasoning is the emerging third pillar, moving beyond correlation to deterministic, model-based root cause identification.

The Bottom Line

Yemini’s framework correctly identifies that we are moving from 'prompt engineering' to 'systems engineering' for AI. If your team is still debugging agents by manually correlating timestamps in dashboards, you are operating with one hand tied behind your back. The era of the 'black box' agent is over; the era of the transparent, causally understood agent has begun.