MemVerdict

Pilot results: 30 of 500 questions per system. Gaps of a few points are within noise; 95% intervals are shown. The full run is in progress.

Compare memory systems

21 head-to-head pages, one per pair, all from the same 30-question pilot of LongMemEval-S (cleaned, 2025-09). Each one-line verdict calls a pair "statistically tied" when the two 95% intervals overlap: at this sample size, most memory systems cannot yet be separated on accuracy.

Cognee comparisons

Cognee results

Hindsight comparisons

Hindsight results

Mem0 comparisons

Mem0 results

LangMem comparisons

LangMem results

Graphiti comparisons

Graphiti results

Plain RAG comparisonsbaseline

Plain RAG results

Full context comparisonsbaseline

Full context results