Full context vs Graphiti: accuracy, cost and latency
Graphiti: our pilot result was withdrawn and is being rerun. This page collects what is known so far: what each system is, how it is set up, and what the vendors report.
Full context
Baseline, not a memory system- Vendor
- Baseline
- Version
- -
- License
- -
- Result
- 83.3% (95% CI 66–93%)
Graphiti
Memory systemRerun in progress- Vendor
- Zep
- Version
- 0.30.2
- License
- Apache-2.0
- Result
- Withdrawn, rerun in progress
Verdict: no head-to-head numbers yet
Pilot result withdrawn: our first adapter added each session as one episode, while Graphiti's own server adds each message as its own episode. Being rerun the canonical way.
Until then, this page shows what each system is, how it is set up, and the scores the vendors report. Those vendor numbers come from different setups, so they are not directly comparable. See the Graphiti page for status.
How the tested systems scored
Full context is highlighted; other systems are muted for context, baselines lightest. Each name links to the system's page.
Accuracy
Higher is betterShare of 30 questions answered correctly (official judge). Whiskers: 95% interval.
- Cognee90.0%74–97%
- Hindsight90.0%74–97%
- Plain RAG86.7%70–95%
- Mem083.3%66–93%
- Full context83.3%66–93%
- LangMem33.3%19–51%
Not shown: Graphiti (pilot result withdrawn, rerun in progress).
Cost per 1,000 questions
Lower is betterUSD billed (prices as of 2026-10-08). Solid: memory side (ingestion and retrieval). Light: answering.
Not shown: Graphiti (pilot result withdrawn, rerun in progress).
Latency per question
Lower is betterRetrieval plus answer, in seconds. Solid: median (p50). Light: up to p90.
- LangMem3.5sp90 5.3s
- Cognee4.0sp90 8.1s
- Mem04.3sp90 7.4s
- Full context4.8sp90 8.0s
- Plain RAG4.8sp90 7.4s
- Hindsight5.4sp90 8.1s
Not shown: Graphiti (pilot result withdrawn, rerun in progress).
Ingestion time
Lower is betterMedian time to load one question's chat history into the system.
Not shown: Graphiti (pilot result withdrawn, rerun in progress).
What each one is
Full context
No memory system. The entire chat history (about 100k tokens) is placed in the reader's prompt.
How we ran it
Official LongMemEval long-context setting with the official answer prompt.
Graphiti
The open-source temporal knowledge-graph engine behind Zep. Facts carry validity intervals, so old facts can be invalidated when they change.
Status of our run
Pilot result withdrawn: our first adapter added each session as one episode, while Graphiti's own server adds each message as its own episode. Being rerun the canonical way.
What the vendors report
Graphiti
| Score | Variant | Reader | Judge | Note | Source |
|---|---|---|---|---|---|
| 71.2% | S | GPT-4o | GPT-4o, official prompts | Zep (paper). | arxiv.org/abs/2501.13956 |
| 90.2% | S | GPT-5.4 | GPT-5.4 with chain-of-thought grading | Zep Cloud. | www.getzep.com/research |
Vendor-reported scores and ours are measured differently, so a gap does not by itself mean either number is wrong. Our pilot uses openai/gpt-6-luna as the reader that writes each answer, gpt-4o-2024-08-06 with the official LongMemEval judge prompt, LongMemEval-S (cleaned, 2025-09) and 30 questions, with each system set up as described above. Scores move with the reader model, the judge model and its prompt, the dataset version, the number of questions, and whether a hosted platform or the open-source package is tested.
Frequently asked questions
Is Full context better than Graphiti?
We do not have a valid result for Graphiti yet. Pilot result withdrawn: our first adapter added each session as one episode, while Graphiti's own server adds each message as its own episode. Being rerun the canonical way. Full context scored 83.3% (95% CI 66–93%) on the same pilot, at $13.1 per 1,000 questions and a median latency of 4.8s.
Which is cheaper, Full context or Graphiti?
Not measured yet: Graphiti's cost will be published with its rerun. Full context: $13.1 per 1,000 questions, all of it answering (no memory-side cost).
Which is faster, Full context or Graphiti?
Not measured yet for Graphiti. Full context answers in a median 4.8s (p90 8.0s) and ingests one chat history in none (median).
How were Full context and Graphiti tested?
Full context was run on the same 30 questions of LongMemEval-S (cleaned, 2025-09) as every other system, by the same harness. For each question, the system ingests that question's chat history, then retrieves context that one reader model (openai/gpt-6-luna) uses to answer with the official LongMemEval prompt. Answers are graded by the official judge (gpt-4o-2024-08-06) and cross-checked by 3 other judges; all judges agreed on 95% of graded answers. Costs are what the providers billed (prices as of 2026-10-08). Full context: Official LongMemEval long-context setting with the official answer prompt. Graphiti: Pilot result withdrawn: our first adapter added each session as one episode, while Graphiti's own server adds each message as its own episode. Being rerun the canonical way.
Which should I use, Full context or Graphiti?
Until the Graphiti rerun is published, our numbers cannot compare them. Vendor-reported scores are listed on this page, but they were measured with different reader models, judges and setups, so they are not directly comparable with our results or with each other. Graphiti is Apache-2.0 licensed.
More comparisons
Other comparisons with Full context
- Cognee vs Full contextStatistically tied on accuracy at n=30; Full context 11× cheaper
- Full context vs HindsightStatistically tied on accuracy at n=30; Full context 13× cheaper
- Full context vs LangMemFull context more accurate and 4.7× cheaper
- Full context vs Mem0Tied on accuracy (83.3% each); Full context 7.1× cheaper
- Full context vs Plain RAGStatistically tied on accuracy at n=30; Plain RAG 2.4× cheaper
Other comparisons with Graphiti
- Cognee vs GraphitiGraphiti: pilot result withdrawn, rerun in progress
- Graphiti vs HindsightGraphiti: pilot result withdrawn, rerun in progress
- Graphiti vs LangMemGraphiti: pilot result withdrawn, rerun in progress
- Graphiti vs Mem0Graphiti: pilot result withdrawn, rerun in progress
- Graphiti vs Plain RAGGraphiti: pilot result withdrawn, rerun in progress