The Harvard MadSys group has released a significant open dataset of agent traces, offering a rare look into real-world LLM agent behavior. Spanning 16 weeks from May 17 to September 5, 2026, the dataset captures 1,213,347 tool calls and 1,186,582 LLM requests across 12,000 sessions. This is production-grade telemetry from the FreeInference gateway, featuring interactions from 14 different agent harnesses, including Claude Code.

Deep Structural Insights

The dataset is structured to support detailed analysis of session dynamics and latency. Each session record includes precise timing metrics like Time To First Token (TTFT) and end-to-end request duration. Crucially, the data includes block-level prefixes tokenized with tiktoken's o200k_base encoder, allowing researchers to replay prefix caches and study how context accumulation impacts performance. The structure also nests subagent sessions directly under the tool calls that spawned them, preserving the hierarchical nature of complex agentic workflows.

Privacy and Anonymization

While the dataset provides extensive metadata, it respects user privacy by excluding actual message text. Sensitive values within tool arguments are replaced with typed placeholders, and rare tool names used by fewer than three accounts are generalized. Providers are anonymized, though the model identifiers remain visible, revealing usage patterns for models like GLM-5.1. This balance allows for rigorous infrastructure analysis without compromising the confidentiality of the 267 distinct user accounts involved.

Integration and Usage

For the builders, the team provides processed Parquet tables alongside the raw JSONL files, making integration with tools like DuckDB seamless. You can query aggregate statistics directly from the Hugging Face Hub without downloading the entire corpus. The dataset is licensed under CC BY 4.0, encouraging widespread adoption in academic and industrial research. This release accompanies the paper 'From Requests to Sessions: A Large-Scale Characterization of Human-Driven Agentic Workloads,' providing the empirical backbone for their findings.

Key Takeaways

  • The dataset covers 16 weeks (May 17–Sept 5, 2026) of production agent traffic.
  • Includes 1,213,347 tool calls and 1,186,582 LLM requests from 12,002 sessions.
  • Supports prefix cache replay via 16-token block IDs tokenized with o200k_base.
  • Released under CC BY 4.0 with processed Parquet tables for easy querying.

The Bottom Line

This is the ground truth data we've been starving for. If you're building agent infrastructure, stop guessing about latency and cache behaviorβ€”download this trace and optimize against reality.