MemVerdict

Pilot results: 30 of 500 questions per system. Gaps of a few points are within noise; 95% intervals are shown. The full run is in progress.

Memory system

Mem0: LongMemEval results

Extracts short memories from each exchange with an LLM, stores them in a vector store, and retrieves them with semantic, keyword and entity signals.

Vendor
Mem0
Version tested
2.2.1 (open source)
License
Apache-2.0

Headline numbers

Pilot: 30 questions of LongMemEval-S (cleaned, 2025-09). With this few questions, gaps of several points between systems are within noise, so read accuracy together with its 95% interval.

Accuracy (official judge)
83.3%
95% CI 66–93% · 25/30 correct
Rank by accuracy
#4 of 6
Among all scored systems, baselines included. Tied with Full context. #3 of 4 memory systems.
Panel accuracy (3-judge majority)
83.3%
Cross-check with gpt-6-sol, claude-sonnet-5.5, gemini-3.1-pro-preview.
Cost per 1,000 questions
$93.0
Memory side $91.6 · answering $1.32 · prices as of 2026-10-08
Latency per question
4.3s p50
p90 7.4s · retrieval plus answer
Ingestion and context
15.2 min
Median time to ingest one chat history; the reader sees a median 50k characters of context.

How Mem0 compares

All scored systems on the same questions, reader and judge. Mem0 is highlighted. Baselines are shown in a lighter grey.

Accuracy

Higher is better

Share of 30 questions answered correctly (official judge). Whiskers: 95% interval.

  1. Cognee
    90.0%74–97%
  2. Hindsight
    90.0%74–97%
  3. Plain RAG
    86.7%70–95%
  4. Mem0
    83.3%66–93%
  5. Full context
    83.3%66–93%
  6. LangMem
    33.3%19–51%

Not shown: Graphiti (pilot result withdrawn, rerun in progress).

Cost per 1,000 questions

Lower is better

USD billed (prices as of 2026-10-08). Solid: memory side (ingestion and retrieval). Light: answering.

  1. Plain RAG
    $5.45
  2. Full context
    $13.1
  3. LangMem
    $61.6
  4. Mem0
    $93.0
  5. Cognee
    $141
  6. Hindsight
    $172

Not shown: Graphiti (pilot result withdrawn, rerun in progress).

Latency per question

Lower is better

Retrieval plus answer, in seconds. Solid: median (p50). Light: up to p90.

  1. LangMem
    3.5sp90 5.3s
  2. Cognee
    4.0sp90 8.1s
  3. Mem0
    4.3sp90 7.4s
  4. Full context
    4.8sp90 8.0s
  5. Plain RAG
    4.8sp90 7.4s
  6. Hindsight
    5.4sp90 8.1s

Not shown: Graphiti (pilot result withdrawn, rerun in progress).

Ingestion time

Lower is better

Median time to load one question's chat history into the system.

  1. Full context
    none
  2. Plain RAG
    1s
  3. Cognee
    2.8 min
  4. LangMem
    8.5 min
  5. Hindsight
    8.8 min
  6. Mem0
    15.2 min

Not shown: Graphiti (pilot result withdrawn, rerun in progress).

Accuracy by question type

LongMemEval groups questions by the memory ability they test. n is the number of pilot questions in each group. Highest: Knowledge update (n=4), Single-session (assistant) (n=3), Preferences (n=2) at 100%. Lowest: Abstention (n=1) at 0%. Category samples are small, so treat these as hints.

  • Single-session (user) n=475% (3/4)
  • Single-session (assistant) n=3100% (3/3)
  • Preferences n=2100% (2/2)
  • Multi-session n=875% (6/8)
  • Knowledge update n=4100% (4/4)
  • Temporal reasoning n=888% (7/8)
  • Abstention n=10% (0/1)

How we ran it

Open-source SDK with local Qdrant, NLP extras installed (spaCy, BM25). Mirrors Mem0's own LongMemEval harness: one add() per user+assistant pair, top_k=200 at search. Historical timestamps are platform-only, so the shared clock simulation supplies dates.

Every system gets each question's chat history, then returns context that the same reader model (openai/gpt-6-luna) uses to answer with the official LongMemEval prompt. Answers are graded by the official judge (gpt-4o-2024-08-06) and cross-checked by 3 other judges. Full details are on the methodology page.

What the vendor reports

LongMemEval scores reported for Mem0
ScoreVariantReaderJudgeNoteSource
83.3%S (cleaned, 2025-09), 30-question pilotopenai/gpt-6-lunagpt-4o-2024-08-06, official promptMemVerdict measurement, Mem0 2.2.1 (open source)This page
94.4%SGPT-5GPT-5 with Mem0's own lenient promptManaged platform, not the open-source SDK.mem0.ai/research

Vendor-reported scores and ours are measured differently, so a gap does not by itself mean either number is wrong. Our pilot uses openai/gpt-6-luna as the reader that writes each answer, gpt-4o-2024-08-06 with the official LongMemEval judge prompt, LongMemEval-S (cleaned, 2025-09) and 30 questions, with Mem0 2.2.1 (open source) set up as described above. Scores move with the reader model, the judge model and its prompt, the dataset version, the number of questions, and whether a hosted platform or the open-source package is tested.

Frequently asked questions

How accurate is Mem0 on LongMemEval?

Mem0 answered 25 of 30 questions correctly (83.3%, 95% CI 66–93%) under the official judge. Rank 4 of 6 scored systems including baselines (tied); 3 of 4 memory systems. The 3-judge panel majority gives 83.3%. This is a 30-question pilot of LongMemEval-S (cleaned, 2025-09); the full 500-question run is in progress.

How much does Mem0 cost to run?

Mem0: $93.0 per 1,000 questions, of which $91.6 (99%) is memory-side (ingestion and retrieval) and $1.32 is answering. Memory-side cost covers ingesting each question's chat history and retrieving from it, as billed by the providers (prices as of 2026-10-08).

How fast is Mem0?

Median latency is 4.3s per question (p90 7.4s), measured as retrieval plus answering. Ingesting one question's chat history takes 15.2 min (median).

Why is this Mem0 score different from the one Mem0 reports?

Mem0 reports 94.4% (LongMemEval-S; reader GPT-5; judge GPT-5 with Mem0's own lenient prompt; Managed platform, not the open-source SDK). We measured 83.3% (95% CI 66–93%). Vendor-reported scores and ours are measured differently, so a gap does not by itself mean either number is wrong. Our pilot uses openai/gpt-6-luna as the reader that writes each answer, gpt-4o-2024-08-06 with the official LongMemEval judge prompt, LongMemEval-S (cleaned, 2025-09) and 30 questions, with Mem0 2.2.1 (open source) set up as described above. Scores move with the reader model, the judge model and its prompt, the dataset version, the number of questions, and whether a hosted platform or the open-source package is tested.

Does Mem0 beat the simple baselines?

Statistically tied on accuracy at n=30: Plain RAG 86.7% (95% CI 70–95%) vs Mem0 83.3% (95% CI 66–93%). The 95% intervals overlap, so the 3.3-point gap could be noise. Mem0 and Full context are tied on accuracy: both answered 25 of 30 questions correctly (83.3%, 95% CI 66–93%).

Compare Mem0 with