If your AI agent is repeating mistakes in fresh sessions, the model is likely not the problem. A new guide published on DEV.to by PLUR argues that developers are wasting cycles swapping LLMs when they should be debugging the memory handoff. The core thesis is that persistence is not a binary state but a three-step pipeline: saving the correction, retrieving it in a new session, and applying it before the agent makes a decision. A successful write proves only the first step, leaving the other two vulnerable to silent failures.

Isolate the Save, Retrieval, and Application

The guide proposes a minimal acceptance test using a fictional project called sample-app. The instruction is simple: "Use the existing test runner; do not introduce a second runner." Developers are told to explicitly ask the agent to save this convention and then inspect the saved record or file. Crucially, a conversational "I'll remember" from the agent is not evidence of a write. The saved artifact must contain the instruction, the project scope, and the source of the decision. Without inspecting the underlying storage, you are debugging blind.

Test in a Genuinely New Session

Once saved, the test requires closing the conversation and starting a new one in the same project. Do not paste the convention into the new prompt, as that tests your prompt engineering, not persistence. Ask the agent to outline a plan for adding a test. Before it edits code, inspect the memory trace or ask the agent to identify the saved context it consulted. This separates retrieval (the record was loaded) from application (the plan respects the convention). A correct answer alone is weak evidence; the agent might infer the convention from the repository code without ever touching memory.

Verify the Integration Layer

The article warns against assuming that a connected MCP memory server guarantees automatic recall. MCP defines how context is exchanged, but it does not prescribe how the host manages that context at runtime. For concrete implementation, the guide points to PLUR’s repository, which distinguishes between tools like plur_learn for storage and plur_recall for retrieval, and runtime adapters that handle automatic injection. It also notes that Anthropic’s context-engineering guidance supports structured notes retrieved later, but this pattern is not a guarantee for any specific setup. You must verify that the context arrives before the relevant decision is made.

Key Takeaways

  • Treat memory failures as three distinct bugs: save, retrieve, or apply. Inspect tool results for each stage separately.
  • Never accept an agent's conversational confirmation as proof of persistence; always verify the saved file or database record.
  • Scope your memory tests to specific projects. Ensure one project's conventions do not leak into another as universal defaults.
  • Test supersession, not just recall. Explicitly replace a convention and verify the agent identifies the current instruction over the old one.

The Bottom Line

Stop treating agent memory like magic. It is a fragile data pipeline. If you cannot inspect the save, retrieval, and application steps independently, you are not running an agent; you are gambling with context windows.