Building error tracking for property management AI agents requires a fundamental shift in how we view state. A recent deep dive from DEV.to argues that admin pages must treat errors as immutable events and resolutions as reversible annotations, rather than mutations of history. This architecture allows operators to answer three critical questions simultaneously: which step delayed a tenant, what the loop cost, and what actually failed. The core decision rule is simple: capture the entire agent loop as one trace, but never overwrite the past.
The Six Signals of Operational Truth
To make this work, every loop must emit six specific signals: correlation_id, step, duration_ms, usage, estimated_cost_micros, and outcome. These aren't just logs; they are the spine of the system. The article emphasizes using monotonic durations to avoid clock skew issues and versioned rate cards for cost estimates so that historical data remains explainable even if pricing changes. Crucially, the outcome field must distinguish between success, timeout, cancellation, and dependency rejection, providing context that generic 'error' flags miss.
Grouping Without Losing Evidence
Grouping errors is necessary for dashboard sanity, but it is inherently lossy. The source material advises normalizing fingerprints by stripping high-cardinality values like request IDs and timestamps, while keeping the original sanitized occurrence separately. This ensures that when an operator needs to reconstruct an incident, the evidence is still there. The admin page should default to unresolved groups, but clicking a row must open a representative event with its full timeline, preserving the search context for further investigation.
Build the Event Path Before the Dashboard
Before worrying about UI charts, you need a robust emission path. The article recommends using OpenTelemetry for vendor-neutral trace propagation and W3C Trace Context for interoperable headers. A Go example is provided to illustrate writing step records to an event stream, emphasizing that telemetry failures must never crash the main application. This involves a bounded, non-recursive fallback path for encoder failures and explicit redaction of sensitive tenant data before serialization.
Resolution Is Not Deletion
A common pitfall in admin tooling is conflating resolution with deletion. The source argues that resolution belongs in an audit record containing the group ID, actor, timestamp, and reason. Reopening a group should add another record, not erase the first. This preserves the evidence chain needed to determine if a failure has returned. The console is only useful if it supports defensible reconstruction, not if it merely looks clean. Mixing resolution with privacy deletion makes the 'resolve' button dangerous and breaks the historical record.
Key Takeaways
- Capture six signals per loop: correlation_id, step, duration_ms, usage, estimated_cost_micros, and outcome.
- Treat errors as immutable events; make resolution a reversible annotation in an audit log.
- Use monotonic durations to avoid clock skew and versioned rate cards for stable cost estimation.
- Test incident reconstruction as an acceptance criterion, verifying that timelines explain latency without double-counting usage.
- Do not use resolution as deletion; keep privacy deletion as a separate, governed operation.
The Bottom Line
If your admin page cannot reconstruct the exact timeline of a failed property workflow without guessing, it is useless for on-call reality. Build for evidence preservation, not dashboard aesthetics.