MemVerdict

Pilot results: 30 of 500 questions per system. Gaps of a few points are within noise; 95% intervals are shown. The full run is in progress.

Hindsight vs Mem0: accuracy, cost and latency

Independent results on the same 30 questions of LongMemEval-S (cleaned, 2025-09), with the same reader model and the official judge. Pilot sample, so accuracy is shown with 95% intervals.

Hindsight

Memory system
Vendor
Vectorize
Version
0.10.3
License
MIT
Result
90.0% (95% CI 74–97%)

Mem0

Memory system
Vendor
Mem0
Version
2.2.1 (open source)
License
Apache-2.0
Result
83.3% (95% CI 66–93%)

Verdict

Generated from the pilot data, n=30 per system
  • AccuracyStatistically tied on accuracy at n=30: Hindsight 90.0% (95% CI 74–97%) vs Mem0 83.3% (95% CI 66–93%). The 95% intervals overlap, so the 6.7-point gap could be noise.
  • CostMem0 costs 1.8× less per 1,000 questions: $93.0 vs $172 for Hindsight.
  • LatencyMem0 answers faster: median 4.3s vs 5.4s for Hindsight per question (p90 7.4s vs 8.1s).
  • IngestionHindsight ingests a chat history faster: 8.8 min vs 15.2 min for Mem0 (median per question).
  • ContextThe reader sees a median of 18k characters of context per question with Hindsight and 50k with Mem0.

Key numbers side by side

Bold marks the better value. Accuracy rows are not bolded when the 95% intervals overlap.

Key numbers for Hindsight and Mem0
MetricHindsightMem0
Accuracy (official judge)90.0%83.3%
95% interval74–97%66–93%
Panel accuracy86.7%83.3%
Cost per 1,000 questions$172$93.0
of which memory side$171$91.6
of which answering$0.71$1.32
Latency p505.4s4.3s
Latency p908.1s7.4s
Ingestion per history (median)8.8 min15.2 min
Context given to reader (median chars)18k50k

Hindsight and Mem0 among all tested systems

Hindsight and Mem0 are highlighted; other systems are muted for context, baselines lightest. Each name links to the system's page.

Accuracy

Higher is better

Share of 30 questions answered correctly (official judge). Whiskers: 95% interval.

  1. Cognee
    90.0%74–97%
  2. Hindsight
    90.0%74–97%
  3. Plain RAG
    86.7%70–95%
  4. Mem0
    83.3%66–93%
  5. Full context
    83.3%66–93%
  6. LangMem
    33.3%19–51%

Not shown: Graphiti (pilot result withdrawn, rerun in progress).

Cost per 1,000 questions

Lower is better

USD billed (prices as of 2026-10-08). Solid: memory side (ingestion and retrieval). Light: answering.

  1. Plain RAG
    $5.45
  2. Full context
    $13.1
  3. LangMem
    $61.6
  4. Mem0
    $93.0
  5. Cognee
    $141
  6. Hindsight
    $172

Not shown: Graphiti (pilot result withdrawn, rerun in progress).

Latency per question

Lower is better

Retrieval plus answer, in seconds. Solid: median (p50). Light: up to p90.

  1. LangMem
    3.5sp90 5.3s
  2. Cognee
    4.0sp90 8.1s
  3. Mem0
    4.3sp90 7.4s
  4. Full context
    4.8sp90 8.0s
  5. Plain RAG
    4.8sp90 7.4s
  6. Hindsight
    5.4sp90 8.1s

Not shown: Graphiti (pilot result withdrawn, rerun in progress).

Ingestion time

Lower is better

Median time to load one question's chat history into the system.

  1. Full context
    none
  2. Plain RAG
    1s
  3. Cognee
    2.8 min
  4. LangMem
    8.5 min
  5. Hindsight
    8.8 min
  6. Mem0
    15.2 min

Not shown: Graphiti (pilot result withdrawn, rerun in progress).

Accuracy by question type

Difference is Hindsight minus Mem0, in percentage points. Each type has only a few pilot questions (n), so one question can move a row by 13 points or more.

Accuracy by question type, Hindsight vs Mem0
Question typenHindsightMem0Difference
Single-session (user)4100%75%+25 pts in favour of Hindsight
Single-session (assistant)333%100%−67 pts in favour of Mem0
Preferences2100%100%0
Multi-session888%75%+13 pts in favour of Hindsight
Knowledge update4100%100%0
Temporal reasoning8100%88%+12 pts in favour of Hindsight
Abstention1100%0%+100 pts in favour of Hindsight

What each one is

Hindsight

Extracts facts into a memory bank on Postgres, consolidates them into observations in the background, and recalls with hybrid search plus a local reranker.

How we ran it

Embedded daemon (pg0 Postgres), one bank per question. Sessions retained with their real timestamp; recall with the question date and default budget. We wait for background consolidation to finish before asking.

All Hindsight resultsSource repository

Mem0

Extracts short memories from each exchange with an LLM, stores them in a vector store, and retrieves them with semantic, keyword and entity signals.

How we ran it

Open-source SDK with local Qdrant, NLP extras installed (spaCy, BM25). Mirrors Mem0's own LongMemEval harness: one add() per user+assistant pair, top_k=200 at search. Historical timestamps are platform-only, so the shared clock simulation supplies dates.

All Mem0 resultsSource repository

What the vendors report

Hindsight

LongMemEval scores reported for Hindsight
ScoreVariantReaderJudgeNoteSource
90.0%S (cleaned, 2025-09), 30-question pilotopenai/gpt-6-lunagpt-4o-2024-08-06, official promptMemVerdict measurement, Hindsight 0.10.3This page
91.4%SGemini 3 ProGPT-OSS-120B-arxiv.org/abs/2512.12818
94.6%S (cleaned)Gemini 3.1 ProGemini 2.5 Flash-Lite-benchmarks.hindsight.vectorize.io

Mem0

LongMemEval scores reported for Mem0
ScoreVariantReaderJudgeNoteSource
83.3%S (cleaned, 2025-09), 30-question pilotopenai/gpt-6-lunagpt-4o-2024-08-06, official promptMemVerdict measurement, Mem0 2.2.1 (open source)This page
94.4%SGPT-5GPT-5 with Mem0's own lenient promptManaged platform, not the open-source SDK.mem0.ai/research

Vendor-reported scores and ours are measured differently, so a gap does not by itself mean either number is wrong. Our pilot uses openai/gpt-6-luna as the reader that writes each answer, gpt-4o-2024-08-06 with the official LongMemEval judge prompt, LongMemEval-S (cleaned, 2025-09) and 30 questions, with each system set up as described above. Scores move with the reader model, the judge model and its prompt, the dataset version, the number of questions, and whether a hosted platform or the open-source package is tested.

Frequently asked questions

Is Hindsight better than Mem0?

Statistically tied on accuracy at n=30: Hindsight 90.0% (95% CI 74–97%) vs Mem0 83.3% (95% CI 66–93%). The 95% intervals overlap, so the 6.7-point gap could be noise. With the 3-judge panel majority instead of the official judge: Hindsight 86.7%, Mem0 83.3%. The largest gap by question type is Single-session (assistant): 1 of 3 vs 3 of 3 correct, in Mem0's favour; each question type has only 1 to 8 questions in this pilot. These are pilot numbers (30 questions each); the full 500-question run will narrow the intervals.

Which is cheaper, Hindsight or Mem0?

Mem0 costs 1.8× less per 1,000 questions: $93.0 vs $172 for Hindsight. Hindsight: $172 per 1,000 questions, of which $171 (over 99%) is memory-side (ingestion and retrieval) and $0.71 is answering. Mem0: $93.0 per 1,000 questions, of which $91.6 (99%) is memory-side (ingestion and retrieval) and $1.32 is answering. Prices as of 2026-10-08.

Which is faster, Hindsight or Mem0?

Mem0 answers faster: median 4.3s vs 5.4s for Hindsight per question (p90 7.4s vs 8.1s). Hindsight ingests a chat history faster: 8.8 min vs 15.2 min for Mem0 (median per question). Latency is measured per question as retrieval plus answering; ingestion is the time to load one question's chat history.

How were Hindsight and Mem0 tested?

Both were run on the same 30 questions of LongMemEval-S (cleaned, 2025-09) as every other system, by the same harness. For each question, the system ingests that question's chat history, then retrieves context that one reader model (openai/gpt-6-luna) uses to answer with the official LongMemEval prompt. Answers are graded by the official judge (gpt-4o-2024-08-06) and cross-checked by 3 other judges; all judges agreed on 95% of graded answers. Costs are what the providers billed (prices as of 2026-10-08). Hindsight: Embedded daemon (pg0 Postgres), one bank per question. Sessions retained with their real timestamp; recall with the question date and default budget. We wait for background consolidation to finish before asking. Mem0: Open-source SDK with local Qdrant, NLP extras installed (spaCy, BM25). Mirrors Mem0's own LongMemEval harness: one add() per user+assistant pair, top_k=200 at search. Historical timestamps are platform-only, so the shared clock simulation supplies dates.

Which should I use, Hindsight or Mem0?

This pilot does not separate them on accuracy, so decide on the other axes and on fit with your stack. Mem0 costs 1.8× less per 1,000 questions: $93.0 vs $172 for Hindsight. Mem0 answers faster: median 4.3s vs 5.4s for Hindsight per question (p90 7.4s vs 8.1s). Hindsight ingests a chat history faster: 8.8 min vs 15.2 min for Mem0 (median per question). Hindsight is MIT and Mem0 is Apache-2.0 licensed. These are pilot numbers (30 questions each); the full 500-question run will narrow the intervals.

More comparisons

All comparisons