Which AI memory actually remembers?
MemVerdict tests long-term memory systems for LLM agents under one set of rules: the same conversations, the same answer model, the same judge. We report accuracy next to what each system costs and how long it takes, and anyone can rerun the numbers.
What we measure
Accuracy
Share of questions answered correctly after the system has stored a long conversation history. Higher is better.
Cost
US dollars per 1,000 questions, counting memory writes, retrieval, re-ranking and the answer call. Lower is better.
Latency
Median time from question to answer, with the memory lookup included. Lower is better.
How the verdict is reached
- Same conditions for every system. One answer model, one memory budget, the same conversations fed in the same order.
- Open-source builds, default settings. Hosted tiers are tested separately and labelled as such.
- Every cost counted. Each model call is billed at list price on the run date, including calls the memory system makes on its own.
- The judge is disclosed and varied. Headline scores use one fixed judge. Two other judges show how much the score moves with the grader.
- Vendors can reply. Maintainers see their results before publication and can propose a better configuration, which we run under the same rules.
Benchmarks
| Benchmark | Role |
|---|---|
| LongMemEval (M) | Main score. Long chat histories with single-session, multi-session, temporal, knowledge-update and abstention questions. |
| MemoryAgentBench | Conflict resolution: does the system notice when a fact has changed? |
| LoCoMo | Reported for reference, since many vendors quote it. |
Systems in the first round
Mem0, Zep (Graphiti), Letta, Cognee, Supermemory, Hindsight, LangMem and MemOS, plus two baselines every memory system should beat: the full conversation in context, and plain retrieval over chunks.
The list may change if a system cannot be run reproducibly. Every exclusion will be explained.
Build a memory system?
The harness and adapters will be open source. Once it is public you can add your system with a pull request; we run it under the same rules and publish the result, whatever it is.