Cognee vs Graphiti: accuracy, cost and latency
Graphiti: our pilot result was withdrawn and is being rerun. This page collects what is known so far: what each system is, how it is set up, and what the vendors report.
Verdict: no head-to-head numbers yet
Pilot result withdrawn: our first adapter added each session as one episode, while Graphiti's own server adds each message as its own episode. Being rerun the canonical way.
Until then, this page shows what each system is, how it is set up, and the scores the vendors report. Those vendor numbers come from different setups, so they are not directly comparable. See the Graphiti page for status.
How the tested systems scored
Cognee is highlighted; other systems are muted for context, baselines lightest. Each name links to the system's page.
Accuracy
Higher is betterShare of 30 questions answered correctly (official judge). Whiskers: 95% interval.
- Cognee90.0%74–97%
- Hindsight90.0%74–97%
- Plain RAG86.7%70–95%
- Mem083.3%66–93%
- Full context83.3%66–93%
- LangMem33.3%19–51%
Not shown: Graphiti (pilot result withdrawn, rerun in progress).
Cost per 1,000 questions
Lower is betterUSD billed (prices as of 2026-10-08). Solid: memory side (ingestion and retrieval). Light: answering.
Not shown: Graphiti (pilot result withdrawn, rerun in progress).
Latency per question
Lower is betterRetrieval plus answer, in seconds. Solid: median (p50). Light: up to p90.
- LangMem3.5sp90 5.3s
- Cognee4.0sp90 8.1s
- Mem04.3sp90 7.4s
- Full context4.8sp90 8.0s
- Plain RAG4.8sp90 7.4s
- Hindsight5.4sp90 8.1s
Not shown: Graphiti (pilot result withdrawn, rerun in progress).
Ingestion time
Lower is betterMedian time to load one question's chat history into the system.
Not shown: Graphiti (pilot result withdrawn, rerun in progress).
What each one is
Cognee
Builds a knowledge graph plus vector index from documents and conversations, then retrieves graph context for a query.
How we ran it
Python SDK in-process with embedded defaults. One text document per session with the session date at the top, then cognify(). Search: Cognee's default HYBRID_COMPLETION with only_context=True and default top_k.
Graphiti
The open-source temporal knowledge-graph engine behind Zep. Facts carry validity intervals, so old facts can be invalidated when they change.
Status of our run
Pilot result withdrawn: our first adapter added each session as one episode, while Graphiti's own server adds each message as its own episode. Being rerun the canonical way.
What the vendors report
Cognee
| Score | Variant | Reader | Judge | Note | Source |
|---|---|---|---|---|---|
| 90.0% | S (cleaned, 2025-09), 30-question pilot | openai/gpt-6-luna | gpt-4o-2024-08-06, official prompt | MemVerdict measurement, Cognee 1.6.3 | This page |
| None found | - | - | - | No self-reported LongMemEval score found. | - |
Graphiti
| Score | Variant | Reader | Judge | Note | Source |
|---|---|---|---|---|---|
| 71.2% | S | GPT-4o | GPT-4o, official prompts | Zep (paper). | arxiv.org/abs/2501.13956 |
| 90.2% | S | GPT-5.4 | GPT-5.4 with chain-of-thought grading | Zep Cloud. | www.getzep.com/research |
Vendor-reported scores and ours are measured differently, so a gap does not by itself mean either number is wrong. Our pilot uses openai/gpt-6-luna as the reader that writes each answer, gpt-4o-2024-08-06 with the official LongMemEval judge prompt, LongMemEval-S (cleaned, 2025-09) and 30 questions, with each system set up as described above. Scores move with the reader model, the judge model and its prompt, the dataset version, the number of questions, and whether a hosted platform or the open-source package is tested.
Frequently asked questions
Is Cognee better than Graphiti?
We do not have a valid result for Graphiti yet. Pilot result withdrawn: our first adapter added each session as one episode, while Graphiti's own server adds each message as its own episode. Being rerun the canonical way. Cognee scored 90.0% (95% CI 74–97%) on the same pilot, at $141 per 1,000 questions and a median latency of 4.0s.
Which is cheaper, Cognee or Graphiti?
Not measured yet: Graphiti's cost will be published with its rerun. Cognee: $141 per 1,000 questions, of which $139 (98%) is memory-side (ingestion and retrieval) and $2.23 is answering.
Which is faster, Cognee or Graphiti?
Not measured yet for Graphiti. Cognee answers in a median 4.0s (p90 8.1s) and ingests one chat history in 2.8 min (median).
How were Cognee and Graphiti tested?
Cognee was run on the same 30 questions of LongMemEval-S (cleaned, 2025-09) as every other system, by the same harness. For each question, the system ingests that question's chat history, then retrieves context that one reader model (openai/gpt-6-luna) uses to answer with the official LongMemEval prompt. Answers are graded by the official judge (gpt-4o-2024-08-06) and cross-checked by 3 other judges; all judges agreed on 95% of graded answers. Costs are what the providers billed (prices as of 2026-10-08). Cognee: Python SDK in-process with embedded defaults. One text document per session with the session date at the top, then cognify(). Search: Cognee's default HYBRID_COMPLETION with only_context=True and default top_k. Graphiti: Pilot result withdrawn: our first adapter added each session as one episode, while Graphiti's own server adds each message as its own episode. Being rerun the canonical way.
Which should I use, Cognee or Graphiti?
Until the Graphiti rerun is published, our numbers cannot compare them. Vendor-reported scores are listed on this page, but they were measured with different reader models, judges and setups, so they are not directly comparable with our results or with each other. Both are Apache-2.0 licensed.
More comparisons
Other comparisons with Cognee
- Cognee vs Full contextStatistically tied on accuracy at n=30; Full context 11× cheaper
- Cognee vs HindsightTied on accuracy (90.0% each); Cognee 1.2× cheaper
- Cognee vs LangMemCognee more accurate; LangMem 2.3× cheaper
- Cognee vs Mem0Statistically tied on accuracy at n=30; Mem0 1.5× cheaper
- Cognee vs Plain RAGStatistically tied on accuracy at n=30; Plain RAG 26× cheaper
Other comparisons with Graphiti
- Full context vs GraphitiGraphiti: pilot result withdrawn, rerun in progress
- Graphiti vs HindsightGraphiti: pilot result withdrawn, rerun in progress
- Graphiti vs LangMemGraphiti: pilot result withdrawn, rerun in progress
- Graphiti vs Mem0Graphiti: pilot result withdrawn, rerun in progress
- Graphiti vs Plain RAGGraphiti: pilot result withdrawn, rerun in progress