MemVerdict

Pilot results: 30 of 500 questions per system. Gaps of a few points are within noise; 95% intervals are shown. The full run is in progress.

LangMem vs Plain RAG: accuracy, cost and latency

A memory system against a no-memory baseline: does the memory layer earn its cost? Independent results on the same 30 questions of LongMemEval-S (cleaned, 2025-09), with the same reader model and the official judge. Pilot sample, so accuracy is shown with 95% intervals.

LangMem

Memory system
Vendor
LangChain
Version
0.0.30
License
MIT
Result
33.3% (95% CI 19–51%)

Plain RAG

Baseline, not a memory system
Vendor
Baseline
Version
-
License
-
Result
86.7% (95% CI 70–95%)

Verdict

Generated from the pilot data, n=30 per system

Is a memory layer worth it here?

Not in this pilot: the Plain RAG baseline is more accurate than LangMem (LangMem 33.3% vs Plain RAG 86.7%; the 95% intervals do not overlap), and LangMem costs 11× more per 1,000 questions. The reader sees a median of 23k characters of context per question with LangMem and 127k with Plain RAG.

  • AccuracyPlain RAG is more accurate: 86.7% (95% CI 70–95%) vs 33.3% (95% CI 19–51%) for LangMem. The 95% intervals do not overlap, even at n=30.
  • CostPlain RAG costs 11× less per 1,000 questions: $5.45 vs $61.6 for LangMem.
  • LatencyLangMem answers faster: median 3.5s vs 4.8s for Plain RAG per question (p90 5.3s vs 7.4s).
  • IngestionPlain RAG ingests a chat history faster: 1s vs 8.5 min for LangMem (median per question).

Key numbers side by side

Bold marks the better value. Accuracy rows are not bolded when the 95% intervals overlap.

Key numbers for LangMem and Plain RAG
MetricLangMemPlain RAG
Accuracy (official judge)33.3%86.7%
95% interval19–51%70–95%
Panel accuracy26.7%86.7%
Cost per 1,000 questions$61.6$5.45
of which memory side$61.0$2.08
of which answering$0.61$3.37
Latency p503.5s4.8s
Latency p905.3s7.4s
Ingestion per history (median)8.5 min1s
Context given to reader (median chars)23k127k

LangMem and Plain RAG among all tested systems

LangMem and Plain RAG are highlighted; other systems are muted for context, baselines lightest. Each name links to the system's page.

Accuracy

Higher is better

Share of 30 questions answered correctly (official judge). Whiskers: 95% interval.

  1. Cognee
    90.0%74–97%
  2. Hindsight
    90.0%74–97%
  3. Plain RAG
    86.7%70–95%
  4. Mem0
    83.3%66–93%
  5. Full context
    83.3%66–93%
  6. LangMem
    33.3%19–51%

Not shown: Graphiti (pilot result withdrawn, rerun in progress).

Cost per 1,000 questions

Lower is better

USD billed (prices as of 2026-10-08). Solid: memory side (ingestion and retrieval). Light: answering.

  1. Plain RAG
    $5.45
  2. Full context
    $13.1
  3. LangMem
    $61.6
  4. Mem0
    $93.0
  5. Cognee
    $141
  6. Hindsight
    $172

Not shown: Graphiti (pilot result withdrawn, rerun in progress).

Latency per question

Lower is better

Retrieval plus answer, in seconds. Solid: median (p50). Light: up to p90.

  1. LangMem
    3.5sp90 5.3s
  2. Cognee
    4.0sp90 8.1s
  3. Mem0
    4.3sp90 7.4s
  4. Full context
    4.8sp90 8.0s
  5. Plain RAG
    4.8sp90 7.4s
  6. Hindsight
    5.4sp90 8.1s

Not shown: Graphiti (pilot result withdrawn, rerun in progress).

Ingestion time

Lower is better

Median time to load one question's chat history into the system.

  1. Full context
    none
  2. Plain RAG
    1s
  3. Cognee
    2.8 min
  4. LangMem
    8.5 min
  5. Hindsight
    8.8 min
  6. Mem0
    15.2 min

Not shown: Graphiti (pilot result withdrawn, rerun in progress).

Accuracy by question type

Difference is LangMem minus Plain RAG, in percentage points. Each type has only a few pilot questions (n), so one question can move a row by 13 points or more.

Accuracy by question type, LangMem vs Plain RAG
Question typenLangMemPlain RAGDifference
Single-session (user)475%100%−25 pts in favour of Plain RAG
Single-session (assistant)30%100%−100 pts in favour of Plain RAG
Preferences250%100%−50 pts in favour of Plain RAG
Multi-session813%75%−62 pts in favour of Plain RAG
Knowledge update450%100%−50 pts in favour of Plain RAG
Temporal reasoning838%88%−50 pts in favour of Plain RAG
Abstention10%0%0

What each one is

LangMem

A memory manager that asks an LLM to extract, consolidate and update memories in a LangGraph store.

How we ran it

create_memory_store_manager with default instructions, one invoke() per session, LangGraph InMemoryStore with OpenAI embeddings, store.search() with the default limit (10).

All LangMem resultsSource repository

Plain RAG

No memory system. Each past session is embedded once; the 10 most similar sessions are given to the reader in date order.

How we ran it

text-embedding-3-small, cosine similarity, top 10 sessions.

All Plain RAG results

What the vendors report

LangMem

LongMemEval scores reported for LangMem
ScoreVariantReaderJudgeNoteSource
33.3%S (cleaned, 2025-09), 30-question pilotopenai/gpt-6-lunagpt-4o-2024-08-06, official promptMemVerdict measurement, LangMem 0.0.30This page
None found---No self-reported LongMemEval score found.-

Vendor-reported scores and ours are measured differently, so a gap does not by itself mean either number is wrong. Our pilot uses openai/gpt-6-luna as the reader that writes each answer, gpt-4o-2024-08-06 with the official LongMemEval judge prompt, LongMemEval-S (cleaned, 2025-09) and 30 questions, with each system set up as described above. Scores move with the reader model, the judge model and its prompt, the dataset version, the number of questions, and whether a hosted platform or the open-source package is tested.

Frequently asked questions

Is LangMem better than Plain RAG?

Plain RAG is more accurate: 86.7% (95% CI 70–95%) vs 33.3% (95% CI 19–51%) for LangMem. The 95% intervals do not overlap, even at n=30. With the 3-judge panel majority instead of the official judge: LangMem 26.7%, Plain RAG 86.7%. The largest gap by question type is Multi-session: 1 of 8 vs 6 of 8 correct, in Plain RAG's favour; each question type has only 1 to 8 questions in this pilot. These are pilot numbers (30 questions each); the full 500-question run will narrow the intervals.

Which is cheaper, LangMem or Plain RAG?

Plain RAG costs 11× less per 1,000 questions: $5.45 vs $61.6 for LangMem. LangMem: $61.6 per 1,000 questions, of which $61.0 (99%) is memory-side (ingestion and retrieval) and $0.61 is answering. Plain RAG: $5.45 per 1,000 questions, of which $2.08 (38%) is memory-side (ingestion and retrieval) and $3.37 is answering. Prices as of 2026-10-08.

Which is faster, LangMem or Plain RAG?

LangMem answers faster: median 3.5s vs 4.8s for Plain RAG per question (p90 5.3s vs 7.4s). Plain RAG ingests a chat history faster: 1s vs 8.5 min for LangMem (median per question). Latency is measured per question as retrieval plus answering; ingestion is the time to load one question's chat history.

How were LangMem and Plain RAG tested?

Both were run on the same 30 questions of LongMemEval-S (cleaned, 2025-09) as every other system, by the same harness. For each question, the system ingests that question's chat history, then retrieves context that one reader model (openai/gpt-6-luna) uses to answer with the official LongMemEval prompt. Answers are graded by the official judge (gpt-4o-2024-08-06) and cross-checked by 3 other judges; all judges agreed on 95% of graded answers. Costs are what the providers billed (prices as of 2026-10-08). LangMem: create_memory_store_manager with default instructions, one invoke() per session, LangGraph InMemoryStore with OpenAI embeddings, store.search() with the default limit (10). Plain RAG: text-embedding-3-small, cosine similarity, top 10 sessions.

Which should I use, LangMem or Plain RAG?

Not in this pilot: the Plain RAG baseline is more accurate than LangMem (LangMem 33.3% vs Plain RAG 86.7%; the 95% intervals do not overlap), and LangMem costs 11× more per 1,000 questions. The reader sees a median of 23k characters of context per question with LangMem and 127k with Plain RAG. If the Plain RAG baseline is as accurate on your own data, it is the simpler option to run; test both on a sample of your real conversations before committing. These are pilot numbers (30 questions each); the full 500-question run will narrow the intervals.

More comparisons

All comparisons