MemVerdict

Pilot results: 30 of 500 questions per system. Gaps of a few points are within noise; 95% intervals are shown. The full run is in progress.

Cognee vs Full context: accuracy, cost and latency

A memory system against a no-memory baseline: does the memory layer earn its cost? Independent results on the same 30 questions of LongMemEval-S (cleaned, 2025-09), with the same reader model and the official judge. Pilot sample, so accuracy is shown with 95% intervals.

Cognee

Memory system
Vendor
Cognee
Version
1.6.3
License
Apache-2.0
Result
90.0% (95% CI 74–97%)

Full context

Baseline, not a memory system
Vendor
Baseline
Version
-
License
-
Result
83.3% (95% CI 66–93%)

Verdict

Generated from the pilot data, n=30 per system

Is a memory layer worth it here?

Not demonstrated in this pilot: Cognee and the Full context baseline are statistically tied on accuracy at n=30 (Cognee 90.0% vs Full context 83.3%; the 95% intervals overlap), and Cognee costs 11× more per 1,000 questions. The reader sees a median of 83k characters of context per question with Cognee and 500k with Full context.

  • AccuracyStatistically tied on accuracy at n=30: Cognee 90.0% (95% CI 74–97%) vs Full context 83.3% (95% CI 66–93%). The 95% intervals overlap, so the 6.7-point gap could be noise.
  • CostFull context costs 11× less per 1,000 questions: $13.1 vs $141 for Cognee.
  • LatencyCognee answers faster: median 4.0s vs 4.8s for Full context per question (p90 8.1s vs 8.0s).
  • IngestionThe Full context baseline needs no ingestion step; Cognee takes 2.8 min to ingest one chat history (median).

Key numbers side by side

Bold marks the better value. Accuracy rows are not bolded when the 95% intervals overlap.

Key numbers for Cognee and Full context
MetricCogneeFull context
Accuracy (official judge)90.0%83.3%
95% interval74–97%66–93%
Panel accuracy90.0%86.7%
Cost per 1,000 questions$141$13.1
of which memory side$139$0.00
of which answering$2.23$13.1
Latency p504.0s4.8s
Latency p908.1s8.0s
Ingestion per history (median)2.8 minnone
Context given to reader (median chars)83k500k

Cognee and Full context among all tested systems

Cognee and Full context are highlighted; other systems are muted for context, baselines lightest. Each name links to the system's page.

Accuracy

Higher is better

Share of 30 questions answered correctly (official judge). Whiskers: 95% interval.

  1. Cognee
    90.0%74–97%
  2. Hindsight
    90.0%74–97%
  3. Plain RAG
    86.7%70–95%
  4. Mem0
    83.3%66–93%
  5. Full context
    83.3%66–93%
  6. LangMem
    33.3%19–51%

Not shown: Graphiti (pilot result withdrawn, rerun in progress).

Cost per 1,000 questions

Lower is better

USD billed (prices as of 2026-10-08). Solid: memory side (ingestion and retrieval). Light: answering.

  1. Plain RAG
    $5.45
  2. Full context
    $13.1
  3. LangMem
    $61.6
  4. Mem0
    $93.0
  5. Cognee
    $141
  6. Hindsight
    $172

Not shown: Graphiti (pilot result withdrawn, rerun in progress).

Latency per question

Lower is better

Retrieval plus answer, in seconds. Solid: median (p50). Light: up to p90.

  1. LangMem
    3.5sp90 5.3s
  2. Cognee
    4.0sp90 8.1s
  3. Mem0
    4.3sp90 7.4s
  4. Full context
    4.8sp90 8.0s
  5. Plain RAG
    4.8sp90 7.4s
  6. Hindsight
    5.4sp90 8.1s

Not shown: Graphiti (pilot result withdrawn, rerun in progress).

Ingestion time

Lower is better

Median time to load one question's chat history into the system.

  1. Full context
    none
  2. Plain RAG
    1s
  3. Cognee
    2.8 min
  4. LangMem
    8.5 min
  5. Hindsight
    8.8 min
  6. Mem0
    15.2 min

Not shown: Graphiti (pilot result withdrawn, rerun in progress).

Accuracy by question type

Difference is Cognee minus Full context, in percentage points. Each type has only a few pilot questions (n), so one question can move a row by 13 points or more.

Accuracy by question type, Cognee vs Full context
Question typenCogneeFull contextDifference
Single-session (user)4100%100%0
Single-session (assistant)3100%100%0
Preferences2100%100%0
Multi-session888%75%+13 pts in favour of Cognee
Knowledge update4100%100%0
Temporal reasoning875%75%0
Abstention1100%0%+100 pts in favour of Cognee

What each one is

Cognee

Builds a knowledge graph plus vector index from documents and conversations, then retrieves graph context for a query.

How we ran it

Python SDK in-process with embedded defaults. One text document per session with the session date at the top, then cognify(). Search: Cognee's default HYBRID_COMPLETION with only_context=True and default top_k.

All Cognee resultsSource repository

Full context

No memory system. The entire chat history (about 100k tokens) is placed in the reader's prompt.

How we ran it

Official LongMemEval long-context setting with the official answer prompt.

All Full context results

What the vendors report

Cognee

LongMemEval scores reported for Cognee
ScoreVariantReaderJudgeNoteSource
90.0%S (cleaned, 2025-09), 30-question pilotopenai/gpt-6-lunagpt-4o-2024-08-06, official promptMemVerdict measurement, Cognee 1.6.3This page
None found---No self-reported LongMemEval score found.-

Vendor-reported scores and ours are measured differently, so a gap does not by itself mean either number is wrong. Our pilot uses openai/gpt-6-luna as the reader that writes each answer, gpt-4o-2024-08-06 with the official LongMemEval judge prompt, LongMemEval-S (cleaned, 2025-09) and 30 questions, with each system set up as described above. Scores move with the reader model, the judge model and its prompt, the dataset version, the number of questions, and whether a hosted platform or the open-source package is tested.

Frequently asked questions

Is Cognee better than Full context?

Statistically tied on accuracy at n=30: Cognee 90.0% (95% CI 74–97%) vs Full context 83.3% (95% CI 66–93%). The 95% intervals overlap, so the 6.7-point gap could be noise. With the 3-judge panel majority instead of the official judge: Cognee 90.0%, Full context 86.7%. No question type separates them by more than one question (each question type has only 1 to 8 questions in this pilot). These are pilot numbers (30 questions each); the full 500-question run will narrow the intervals.

Which is cheaper, Cognee or Full context?

Full context costs 11× less per 1,000 questions: $13.1 vs $141 for Cognee. Cognee: $141 per 1,000 questions, of which $139 (98%) is memory-side (ingestion and retrieval) and $2.23 is answering. Full context: $13.1 per 1,000 questions, all of it answering (no memory-side cost). Prices as of 2026-10-08.

Which is faster, Cognee or Full context?

Cognee answers faster: median 4.0s vs 4.8s for Full context per question (p90 8.1s vs 8.0s). The Full context baseline needs no ingestion step; Cognee takes 2.8 min to ingest one chat history (median). Latency is measured per question as retrieval plus answering; ingestion is the time to load one question's chat history.

How were Cognee and Full context tested?

Both were run on the same 30 questions of LongMemEval-S (cleaned, 2025-09) as every other system, by the same harness. For each question, the system ingests that question's chat history, then retrieves context that one reader model (openai/gpt-6-luna) uses to answer with the official LongMemEval prompt. Answers are graded by the official judge (gpt-4o-2024-08-06) and cross-checked by 3 other judges; all judges agreed on 95% of graded answers. Costs are what the providers billed (prices as of 2026-10-08). Cognee: Python SDK in-process with embedded defaults. One text document per session with the session date at the top, then cognify(). Search: Cognee's default HYBRID_COMPLETION with only_context=True and default top_k. Full context: Official LongMemEval long-context setting with the official answer prompt.

Which should I use, Cognee or Full context?

Not demonstrated in this pilot: Cognee and the Full context baseline are statistically tied on accuracy at n=30 (Cognee 90.0% vs Full context 83.3%; the 95% intervals overlap), and Cognee costs 11× more per 1,000 questions. The reader sees a median of 83k characters of context per question with Cognee and 500k with Full context. If the Full context baseline is as accurate on your own data, it is the simpler option to run; test both on a sample of your real conversations before committing. These are pilot numbers (30 questions each); the full 500-question run will narrow the intervals.

More comparisons

All comparisons