past.dev launched its self-serve memory API on September 30, 2026, positioning itself as the first production-ready solution for unlimited, accurate agent memory. The platform distinguishes itself from standard vector RAG by tracking temporal truth: it records what is true now, what it replaced, and exactly when each fact changed. This approach addresses a critical failure mode in current agent architectures, where similarity search often conflates outdated context with current reality.

Dominating the BEAM Benchmark

The core technical claim rests on performance across the BEAM benchmark, published at ICLR 2026. past.dev reports a 92.08% accuracy at 100K tokens, maintaining an 85.03% score even when history expands to 10M tokens. This represents a mere 7.05-point drop across a hundredfold increase in context size. In comparison, the next best published system, Exabase M-1, scores 68.0% at the 10M token mark, while Hindsight and mem0 fall to 64.1% and 48.6% respectively.

How It Works: Temporal Truth Over Vector Similarity

The API operates on a three-call model: ingest, wait, and recall. Unlike vector stores that retrieve text based on semantic similarity, past.dev derives facts, rules, and dated events from ingested sources like emails and tickets. When a user queries the system, it returns ranked documents with verbatim excerpts, dates, and source IDs. For example, if a budget changes from $32k to $40k, the system cites the newer email while retaining the older call note as historical context, ensuring the agent answers with the current truth rather than the most similar-sounding past statement.

Transparency and Known Limitations

The team publishes a full breakdown of abilities, revealing where the system still struggles. While event ordering and summarization remain near-perfect at 100%, contradiction resolution hovers between 73% and 79%, and multi-session reasoning drops sharply to 28.17% at 10M tokens. These figures, derived from a September 29, 2026 run, indicate that while past.dev solves temporal accuracy, complex reasoning across fragmented long-term histories remains an open challenge for the entire industry.

Key Takeaways

  • past.dev offers 150,000 free credits ($45 value) for self-serve sign-ups, with no credit card required initially.
  • The platform is SOC 2 Type II audited and ISO 27001/27701 certified, with customer data never used for model training.
  • Enterprise plans offer dedicated regions and self-hosting options for teams requiring strict data control.
  • Benchmarks are reproducible via the public GitHub harness 'pastdotdev/benchmarks', which includes per-question results.

The Bottom Line

Past.dev finally solves the temporal reasoning problem that has plagued agent memory for years, but the sharp drop in multi-session reasoning at scale proves that true long-term cognitive continuity is still an unsolved frontier.