MemVerdict

Pilot results: 30 of 500 questions per system. Gaps of a few points are within noise; 95% intervals are shown. The full run is in progress.

Baseline, not a memory system

Full context baseline: LongMemEval results

No memory system. The entire chat history (about 100k tokens) is placed in the reader's prompt.

Vendor
Baseline
Version tested
-
License
-

Headline numbers

Pilot: 30 questions of LongMemEval-S (cleaned, 2025-09). With this few questions, gaps of several points between systems are within noise, so read accuracy together with its 95% interval.

Accuracy (official judge)
83.3%
95% CI 66–93% · 25/30 correct
Rank by accuracy
#4 of 6
Among all scored systems, baselines included. Tied with Mem0.
Panel accuracy (3-judge majority)
86.7%
Cross-check with gpt-6-sol, claude-sonnet-5.5, gemini-3.1-pro-preview.
Cost per 1,000 questions
$13.1
Memory side $0.00 · answering $13.1 · prices as of 2026-10-08
Latency per question
4.8s p50
p90 8.0s · retrieval plus answer
Ingestion and context
none
Median time to ingest one chat history; the reader sees a median 500k characters of context.

How Full context compares

All scored systems on the same questions, reader and judge. Full context is highlighted. Baselines are shown in a lighter grey.

Accuracy

Higher is better

Share of 30 questions answered correctly (official judge). Whiskers: 95% interval.

  1. Cognee
    90.0%74–97%
  2. Hindsight
    90.0%74–97%
  3. Plain RAG
    86.7%70–95%
  4. Mem0
    83.3%66–93%
  5. Full context
    83.3%66–93%
  6. LangMem
    33.3%19–51%

Not shown: Graphiti (pilot result withdrawn, rerun in progress).

Cost per 1,000 questions

Lower is better

USD billed (prices as of 2026-10-08). Solid: memory side (ingestion and retrieval). Light: answering.

  1. Plain RAG
    $5.45
  2. Full context
    $13.1
  3. LangMem
    $61.6
  4. Mem0
    $93.0
  5. Cognee
    $141
  6. Hindsight
    $172

Not shown: Graphiti (pilot result withdrawn, rerun in progress).

Latency per question

Lower is better

Retrieval plus answer, in seconds. Solid: median (p50). Light: up to p90.

  1. LangMem
    3.5sp90 5.3s
  2. Cognee
    4.0sp90 8.1s
  3. Mem0
    4.3sp90 7.4s
  4. Full context
    4.8sp90 8.0s
  5. Plain RAG
    4.8sp90 7.4s
  6. Hindsight
    5.4sp90 8.1s

Not shown: Graphiti (pilot result withdrawn, rerun in progress).

Ingestion time

Lower is better

Median time to load one question's chat history into the system.

  1. Full context
    none
  2. Plain RAG
    1s
  3. Cognee
    2.8 min
  4. LangMem
    8.5 min
  5. Hindsight
    8.8 min
  6. Mem0
    15.2 min

Not shown: Graphiti (pilot result withdrawn, rerun in progress).

Accuracy by question type

LongMemEval groups questions by the memory ability they test. n is the number of pilot questions in each group. Highest: Knowledge update (n=4), Single-session (assistant) (n=3), Preferences (n=2), Single-session (user) (n=4) at 100%. Lowest: Abstention (n=1) at 0%. Category samples are small, so treat these as hints.

  • Single-session (user) n=4100% (4/4)
  • Single-session (assistant) n=3100% (3/3)
  • Preferences n=2100% (2/2)
  • Multi-session n=875% (6/8)
  • Knowledge update n=4100% (4/4)
  • Temporal reasoning n=875% (6/8)
  • Abstention n=10% (0/1)

How we ran it

Official LongMemEval long-context setting with the official answer prompt.

Every system gets each question's chat history, then returns context that the same reader model (openai/gpt-6-luna) uses to answer with the official LongMemEval prompt. Answers are graded by the official judge (gpt-4o-2024-08-06) and cross-checked by 3 other judges. Full details are on the methodology page.

Frequently asked questions

What is the Full context baseline?

No memory system. The entire chat history (about 100k tokens) is placed in the reader's prompt. Setup: Official LongMemEval long-context setting with the official answer prompt. We include it as a reference point for the memory systems.

How accurate is Full context on LongMemEval?

Full context answered 25 of 30 questions correctly (83.3%, 95% CI 66–93%) under the official judge. Rank 4 of 6 scored systems including baselines (tied). The 3-judge panel majority gives 86.7%. This is a 30-question pilot of LongMemEval-S (cleaned, 2025-09); the full 500-question run is in progress.

How much does Full context cost to run?

Full context: $13.1 per 1,000 questions, all of it answering (no memory-side cost). Memory-side cost covers ingesting each question's chat history and retrieving from it, as billed by the providers (prices as of 2026-10-08).

How fast is Full context?

Median latency is 4.8s per question (p90 8.0s), measured as retrieval plus answering. It needs no ingestion step.

Why include a Full context baseline?

Baselines show whether a memory layer adds anything over a simple approach on the same questions, reader and judge. In this pilot, no memory system is clearly more accurate than Full context, while Full context is clearly more accurate than LangMem; the rest are statistically tied with it at n=30.

Compare Full context with