Memory systems we have tested
Each system ran the same 30 questions of LongMemEval-S (cleaned, 2025-09), with the same reader model and the official judge (gpt-4o-2024-08-06). These are pilot numbers, so every accuracy comes with its 95% interval. Baselines are included as reference points: they are not memory systems. Compare any two.
Memory systems
Cognee
Cognee · 1.6.3Builds a knowledge graph plus vector index from documents and conversations, then retrieves graph context for a query.
- Accuracy
- 90.0%
- CI 74–97%
- Per 1k questions
- $141
- Rank T-1 of 6
- Median latency
- 4.0s
- p90 8.1s
Hindsight
Vectorize · 0.10.3Extracts facts into a memory bank on Postgres, consolidates them into observations in the background, and recalls with hybrid search plus a local reranker.
- Accuracy
- 90.0%
- CI 74–97%
- Per 1k questions
- $172
- Rank T-1 of 6
- Median latency
- 5.4s
- p90 8.1s
Mem0
Mem0 · 2.2.1 (open source)Extracts short memories from each exchange with an LLM, stores them in a vector store, and retrieves them with semantic, keyword and entity signals.
- Accuracy
- 83.3%
- CI 66–93%
- Per 1k questions
- $93.0
- Rank T-4 of 6
- Median latency
- 4.3s
- p90 7.4s
LangMem
LangChain · 0.0.30A memory manager that asks an LLM to extract, consolidate and update memories in a LangGraph store.
- Accuracy
- 33.3%
- CI 19–51%
- Per 1k questions
- $61.6
- Rank 6 of 6
- Median latency
- 3.5s
- p90 5.3s
Graphiti
Zep · 0.30.2The open-source temporal knowledge-graph engine behind Zep. Facts carry validity intervals, so old facts can be invalidated when they change.
No result yet: the pilot run was withdrawn and is being rerun.
Baselines
Simple approaches with no memory layer. If a memory system cannot beat them, its extra cost and moving parts are hard to justify on this benchmark.
Plain RAG
BaselineNo memory system. Each past session is embedded once; the 10 most similar sessions are given to the reader in date order.
- Accuracy
- 86.7%
- CI 70–95%
- Per 1k questions
- $5.45
- Rank 3 of 6
- Median latency
- 4.8s
- p90 7.4s
Full context
BaselineNo memory system. The entire chat history (about 100k tokens) is placed in the reader's prompt.
- Accuracy
- 83.3%
- CI 66–93%
- Per 1k questions
- $13.1
- Rank T-4 of 6
- Median latency
- 4.8s
- p90 8.0s
Coming soon
EverOS
EverMind · not run yet
MemOS
MemTensor · not run yet
Letta
Letta · not run yet
Supermemory
Supermemory · not run yet