Graphiti vs Plain RAG: accuracy, cost and latency
Graphiti: our pilot result was withdrawn and is being rerun. This page collects what is known so far: what each system is, how it is set up, and what the vendors report.
Verdict: no head-to-head numbers yet
Pilot result withdrawn: our first adapter added each session as one episode, while Graphiti's own server adds each message as its own episode. Being rerun the canonical way.
Until then, this page shows what each system is, how it is set up, and the scores the vendors report. Those vendor numbers come from different setups, so they are not directly comparable. See the Graphiti page for status.
How the tested systems scored
Plain RAG is highlighted; other systems are muted for context, baselines lightest. Each name links to the system's page.
Accuracy
Higher is betterShare of 30 questions answered correctly (official judge). Whiskers: 95% interval.
- Cognee90.0%74–97%
- Hindsight90.0%74–97%
- Plain RAG86.7%70–95%
- Mem083.3%66–93%
- Full context83.3%66–93%
- LangMem33.3%19–51%
Not shown: Graphiti (pilot result withdrawn, rerun in progress).
Cost per 1,000 questions
Lower is betterUSD billed (prices as of 2026-10-08). Solid: memory side (ingestion and retrieval). Light: answering.
Not shown: Graphiti (pilot result withdrawn, rerun in progress).
Latency per question
Lower is betterRetrieval plus answer, in seconds. Solid: median (p50). Light: up to p90.
- LangMem3.5sp90 5.3s
- Cognee4.0sp90 8.1s
- Mem04.3sp90 7.4s
- Full context4.8sp90 8.0s
- Plain RAG4.8sp90 7.4s
- Hindsight5.4sp90 8.1s
Not shown: Graphiti (pilot result withdrawn, rerun in progress).
Ingestion time
Lower is betterMedian time to load one question's chat history into the system.
Not shown: Graphiti (pilot result withdrawn, rerun in progress).
What each one is
Graphiti
The open-source temporal knowledge-graph engine behind Zep. Facts carry validity intervals, so old facts can be invalidated when they change.
Status of our run
Pilot result withdrawn: our first adapter added each session as one episode, while Graphiti's own server adds each message as its own episode. Being rerun the canonical way.
Plain RAG
No memory system. Each past session is embedded once; the 10 most similar sessions are given to the reader in date order.
How we ran it
text-embedding-3-small, cosine similarity, top 10 sessions.
What the vendors report
Graphiti
| Score | Variant | Reader | Judge | Note | Source |
|---|---|---|---|---|---|
| 71.2% | S | GPT-4o | GPT-4o, official prompts | Zep (paper). | arxiv.org/abs/2501.13956 |
| 90.2% | S | GPT-5.4 | GPT-5.4 with chain-of-thought grading | Zep Cloud. | www.getzep.com/research |
Vendor-reported scores and ours are measured differently, so a gap does not by itself mean either number is wrong. Our pilot uses openai/gpt-6-luna as the reader that writes each answer, gpt-4o-2024-08-06 with the official LongMemEval judge prompt, LongMemEval-S (cleaned, 2025-09) and 30 questions, with each system set up as described above. Scores move with the reader model, the judge model and its prompt, the dataset version, the number of questions, and whether a hosted platform or the open-source package is tested.
Frequently asked questions
Is Graphiti better than Plain RAG?
We do not have a valid result for Graphiti yet. Pilot result withdrawn: our first adapter added each session as one episode, while Graphiti's own server adds each message as its own episode. Being rerun the canonical way. Plain RAG scored 86.7% (95% CI 70–95%) on the same pilot, at $5.45 per 1,000 questions and a median latency of 4.8s.
Which is cheaper, Graphiti or Plain RAG?
Not measured yet: Graphiti's cost will be published with its rerun. Plain RAG: $5.45 per 1,000 questions, of which $2.08 (38%) is memory-side (ingestion and retrieval) and $3.37 is answering.
Which is faster, Graphiti or Plain RAG?
Not measured yet for Graphiti. Plain RAG answers in a median 4.8s (p90 7.4s) and ingests one chat history in 1s (median).
How were Graphiti and Plain RAG tested?
Plain RAG was run on the same 30 questions of LongMemEval-S (cleaned, 2025-09) as every other system, by the same harness. For each question, the system ingests that question's chat history, then retrieves context that one reader model (openai/gpt-6-luna) uses to answer with the official LongMemEval prompt. Answers are graded by the official judge (gpt-4o-2024-08-06) and cross-checked by 3 other judges; all judges agreed on 95% of graded answers. Costs are what the providers billed (prices as of 2026-10-08). Graphiti: Pilot result withdrawn: our first adapter added each session as one episode, while Graphiti's own server adds each message as its own episode. Being rerun the canonical way. Plain RAG: text-embedding-3-small, cosine similarity, top 10 sessions.
Which should I use, Graphiti or Plain RAG?
Until the Graphiti rerun is published, our numbers cannot compare them. Vendor-reported scores are listed on this page, but they were measured with different reader models, judges and setups, so they are not directly comparable with our results or with each other. Graphiti is Apache-2.0 licensed.
More comparisons
Other comparisons with Graphiti
- Cognee vs GraphitiGraphiti: pilot result withdrawn, rerun in progress
- Full context vs GraphitiGraphiti: pilot result withdrawn, rerun in progress
- Graphiti vs HindsightGraphiti: pilot result withdrawn, rerun in progress
- Graphiti vs LangMemGraphiti: pilot result withdrawn, rerun in progress
- Graphiti vs Mem0Graphiti: pilot result withdrawn, rerun in progress
Other comparisons with Plain RAG
- Cognee vs Plain RAGStatistically tied on accuracy at n=30; Plain RAG 26× cheaper
- Full context vs Plain RAGStatistically tied on accuracy at n=30; Plain RAG 2.4× cheaper
- Hindsight vs Plain RAGStatistically tied on accuracy at n=30; Plain RAG 31× cheaper
- LangMem vs Plain RAGPlain RAG more accurate and 11× cheaper
- Mem0 vs Plain RAGStatistically tied on accuracy at n=30; Plain RAG 17× cheaper