MemVerdict

Pilot results: 30 of 500 questions per system. Gaps of a few points are within noise; 95% intervals are shown. The full run is in progress.

Graphiti vs LangMem: accuracy, cost and latency

Graphiti: our pilot result was withdrawn and is being rerun. This page collects what is known so far: what each system is, how it is set up, and what the vendors report.

Graphiti

Memory systemRerun in progress
Vendor
Zep
Version
0.30.2
License
Apache-2.0
Result
Withdrawn, rerun in progress

LangMem

Memory system
Vendor
LangChain
Version
0.0.30
License
MIT
Result
33.3% (95% CI 19–51%)

Verdict: no head-to-head numbers yet

Pilot result withdrawn: our first adapter added each session as one episode, while Graphiti's own server adds each message as its own episode. Being rerun the canonical way.

Until then, this page shows what each system is, how it is set up, and the scores the vendors report. Those vendor numbers come from different setups, so they are not directly comparable. See the Graphiti page for status.

How the tested systems scored

LangMem is highlighted; other systems are muted for context, baselines lightest. Each name links to the system's page.

Accuracy

Higher is better

Share of 30 questions answered correctly (official judge). Whiskers: 95% interval.

  1. Cognee
    90.0%74–97%
  2. Hindsight
    90.0%74–97%
  3. Plain RAG
    86.7%70–95%
  4. Mem0
    83.3%66–93%
  5. Full context
    83.3%66–93%
  6. LangMem
    33.3%19–51%

Not shown: Graphiti (pilot result withdrawn, rerun in progress).

Cost per 1,000 questions

Lower is better

USD billed (prices as of 2026-10-08). Solid: memory side (ingestion and retrieval). Light: answering.

  1. Plain RAG
    $5.45
  2. Full context
    $13.1
  3. LangMem
    $61.6
  4. Mem0
    $93.0
  5. Cognee
    $141
  6. Hindsight
    $172

Not shown: Graphiti (pilot result withdrawn, rerun in progress).

Latency per question

Lower is better

Retrieval plus answer, in seconds. Solid: median (p50). Light: up to p90.

  1. LangMem
    3.5sp90 5.3s
  2. Cognee
    4.0sp90 8.1s
  3. Mem0
    4.3sp90 7.4s
  4. Full context
    4.8sp90 8.0s
  5. Plain RAG
    4.8sp90 7.4s
  6. Hindsight
    5.4sp90 8.1s

Not shown: Graphiti (pilot result withdrawn, rerun in progress).

Ingestion time

Lower is better

Median time to load one question's chat history into the system.

  1. Full context
    none
  2. Plain RAG
    1s
  3. Cognee
    2.8 min
  4. LangMem
    8.5 min
  5. Hindsight
    8.8 min
  6. Mem0
    15.2 min

Not shown: Graphiti (pilot result withdrawn, rerun in progress).

What each one is

Graphiti

The open-source temporal knowledge-graph engine behind Zep. Facts carry validity intervals, so old facts can be invalidated when they change.

Status of our run

Pilot result withdrawn: our first adapter added each session as one episode, while Graphiti's own server adds each message as its own episode. Being rerun the canonical way.

All Graphiti resultsSource repository

LangMem

A memory manager that asks an LLM to extract, consolidate and update memories in a LangGraph store.

How we ran it

create_memory_store_manager with default instructions, one invoke() per session, LangGraph InMemoryStore with OpenAI embeddings, store.search() with the default limit (10).

All LangMem resultsSource repository

What the vendors report

Graphiti

LongMemEval scores reported for Graphiti
ScoreVariantReaderJudgeNoteSource
71.2%SGPT-4oGPT-4o, official promptsZep (paper).arxiv.org/abs/2501.13956
90.2%SGPT-5.4GPT-5.4 with chain-of-thought gradingZep Cloud.www.getzep.com/research

LangMem

LongMemEval scores reported for LangMem
ScoreVariantReaderJudgeNoteSource
33.3%S (cleaned, 2025-09), 30-question pilotopenai/gpt-6-lunagpt-4o-2024-08-06, official promptMemVerdict measurement, LangMem 0.0.30This page
None found---No self-reported LongMemEval score found.-

Vendor-reported scores and ours are measured differently, so a gap does not by itself mean either number is wrong. Our pilot uses openai/gpt-6-luna as the reader that writes each answer, gpt-4o-2024-08-06 with the official LongMemEval judge prompt, LongMemEval-S (cleaned, 2025-09) and 30 questions, with each system set up as described above. Scores move with the reader model, the judge model and its prompt, the dataset version, the number of questions, and whether a hosted platform or the open-source package is tested.

Frequently asked questions

Is Graphiti better than LangMem?

We do not have a valid result for Graphiti yet. Pilot result withdrawn: our first adapter added each session as one episode, while Graphiti's own server adds each message as its own episode. Being rerun the canonical way. LangMem scored 33.3% (95% CI 19–51%) on the same pilot, at $61.6 per 1,000 questions and a median latency of 3.5s.

Which is cheaper, Graphiti or LangMem?

Not measured yet: Graphiti's cost will be published with its rerun. LangMem: $61.6 per 1,000 questions, of which $61.0 (99%) is memory-side (ingestion and retrieval) and $0.61 is answering.

Which is faster, Graphiti or LangMem?

Not measured yet for Graphiti. LangMem answers in a median 3.5s (p90 5.3s) and ingests one chat history in 8.5 min (median).

How were Graphiti and LangMem tested?

LangMem was run on the same 30 questions of LongMemEval-S (cleaned, 2025-09) as every other system, by the same harness. For each question, the system ingests that question's chat history, then retrieves context that one reader model (openai/gpt-6-luna) uses to answer with the official LongMemEval prompt. Answers are graded by the official judge (gpt-4o-2024-08-06) and cross-checked by 3 other judges; all judges agreed on 95% of graded answers. Costs are what the providers billed (prices as of 2026-10-08). Graphiti: Pilot result withdrawn: our first adapter added each session as one episode, while Graphiti's own server adds each message as its own episode. Being rerun the canonical way. LangMem: create_memory_store_manager with default instructions, one invoke() per session, LangGraph InMemoryStore with OpenAI embeddings, store.search() with the default limit (10).

Which should I use, Graphiti or LangMem?

Until the Graphiti rerun is published, our numbers cannot compare them. Vendor-reported scores are listed on this page, but they were measured with different reader models, judges and setups, so they are not directly comparable with our results or with each other. Graphiti is Apache-2.0 and LangMem is MIT licensed.

More comparisons

All comparisons