MemVerdict

Pilot results: 30 of 500 questions per system. Gaps of a few points are within noise; 95% intervals are shown. The full run is in progress.

Memory systemRerun in progress

Graphiti: LongMemEval results

The open-source temporal knowledge-graph engine behind Zep. Facts carry validity intervals, so old facts can be invalidated when they change.

Vendor
Zep
Version tested
0.30.2
License
Apache-2.0

No result yet: rerun in progress

Pilot result withdrawn: our first adapter added each session as one episode, while Graphiti's own server adds each message as its own episode. Being rerun the canonical way.

We publish no number for Graphiti until the rerun is complete. The vendor-reported scores below were measured by Zep with its own setups.

How the other systems scored

All scored systems on the same questions, reader and judge. Baselines are shown in a lighter grey.

Accuracy

Higher is better

Share of 30 questions answered correctly (official judge). Whiskers: 95% interval.

  1. Cognee
    90.0%74–97%
  2. Hindsight
    90.0%74–97%
  3. Plain RAG
    86.7%70–95%
  4. Mem0
    83.3%66–93%
  5. Full context
    83.3%66–93%
  6. LangMem
    33.3%19–51%

Not shown: Graphiti (pilot result withdrawn, rerun in progress).

Cost per 1,000 questions

Lower is better

USD billed (prices as of 2026-10-08). Solid: memory side (ingestion and retrieval). Light: answering.

  1. Plain RAG
    $5.45
  2. Full context
    $13.1
  3. LangMem
    $61.6
  4. Mem0
    $93.0
  5. Cognee
    $141
  6. Hindsight
    $172

Not shown: Graphiti (pilot result withdrawn, rerun in progress).

Latency per question

Lower is better

Retrieval plus answer, in seconds. Solid: median (p50). Light: up to p90.

  1. LangMem
    3.5sp90 5.3s
  2. Cognee
    4.0sp90 8.1s
  3. Mem0
    4.3sp90 7.4s
  4. Full context
    4.8sp90 8.0s
  5. Plain RAG
    4.8sp90 7.4s
  6. Hindsight
    5.4sp90 8.1s

Not shown: Graphiti (pilot result withdrawn, rerun in progress).

Ingestion time

Lower is better

Median time to load one question's chat history into the system.

  1. Full context
    none
  2. Plain RAG
    1s
  3. Cognee
    2.8 min
  4. LangMem
    8.5 min
  5. Hindsight
    8.8 min
  6. Mem0
    15.2 min

Not shown: Graphiti (pilot result withdrawn, rerun in progress).

How we ran it

The first pilot run of Graphiti was withdrawn (see the status above); the rerun uses the same harness as every other system.

Every system gets each question's chat history, then returns context that the same reader model uses to answer with the official LongMemEval prompt. Answers are graded by the official judge (gpt-4o-2024-08-06) and cross-checked by 3 other judges. Full details are on the methodology page.

What the vendor reports

LongMemEval scores reported for Graphiti
ScoreVariantReaderJudgeNoteSource
71.2%SGPT-4oGPT-4o, official promptsZep (paper).arxiv.org/abs/2501.13956
90.2%SGPT-5.4GPT-5.4 with chain-of-thought gradingZep Cloud.www.getzep.com/research

Vendor-reported scores and ours are measured differently, so a gap does not by itself mean either number is wrong. Our pilot, and the rerun, use openai/gpt-6-luna as the reader that writes each answer, gpt-4o-2024-08-06 with the official LongMemEval judge prompt, LongMemEval-S (cleaned, 2025-09) and 30 questions, with Graphiti 0.30.2. Scores move with the reader model, the judge model and its prompt, the dataset version, the number of questions, and whether a hosted platform or the open-source package is tested.

Frequently asked questions

What is Graphiti?

The open-source temporal knowledge-graph engine behind Zep. Facts carry validity intervals, so old facts can be invalidated when they change. It is developed by Zep and released under the Apache-2.0 license.

Why is there no Graphiti score on MemVerdict yet?

Pilot result withdrawn: our first adapter added each session as one episode, while Graphiti's own server adds each message as its own episode. Being rerun the canonical way.

What LongMemEval score does Zep report for Graphiti?

Reported: 71.2% (LongMemEval-S; reader GPT-4o; judge GPT-4o, official prompts; Zep (paper)) and 90.2% (LongMemEval-S; reader GPT-5.4; judge GPT-5.4 with chain-of-thought grading; Zep Cloud). These were measured with setups different from ours, so they are not directly comparable.

When will Graphiti results be published?

When the rerun finishes, this page and every comparison involving Graphiti will show its accuracy, cost and latency from the same harness, reader and judge as the other systems.

Compare Graphiti with