Graphiti: LongMemEval results
The open-source temporal knowledge-graph engine behind Zep. Facts carry validity intervals, so old facts can be invalidated when they change.
- Vendor
- Zep
- Version tested
- 0.30.2
- License
- Apache-2.0
- Repository
- github.com/getzep/graphiti
No result yet: rerun in progress
Pilot result withdrawn: our first adapter added each session as one episode, while Graphiti's own server adds each message as its own episode. Being rerun the canonical way.
We publish no number for Graphiti until the rerun is complete. The vendor-reported scores below were measured by Zep with its own setups.
How the other systems scored
All scored systems on the same questions, reader and judge. Baselines are shown in a lighter grey.
Accuracy
Higher is betterShare of 30 questions answered correctly (official judge). Whiskers: 95% interval.
- Cognee90.0%74–97%
- Hindsight90.0%74–97%
- Plain RAG86.7%70–95%
- Mem083.3%66–93%
- Full context83.3%66–93%
- LangMem33.3%19–51%
Not shown: Graphiti (pilot result withdrawn, rerun in progress).
Cost per 1,000 questions
Lower is betterUSD billed (prices as of 2026-10-08). Solid: memory side (ingestion and retrieval). Light: answering.
Not shown: Graphiti (pilot result withdrawn, rerun in progress).
Latency per question
Lower is betterRetrieval plus answer, in seconds. Solid: median (p50). Light: up to p90.
- LangMem3.5sp90 5.3s
- Cognee4.0sp90 8.1s
- Mem04.3sp90 7.4s
- Full context4.8sp90 8.0s
- Plain RAG4.8sp90 7.4s
- Hindsight5.4sp90 8.1s
Not shown: Graphiti (pilot result withdrawn, rerun in progress).
Ingestion time
Lower is betterMedian time to load one question's chat history into the system.
Not shown: Graphiti (pilot result withdrawn, rerun in progress).
How we ran it
The first pilot run of Graphiti was withdrawn (see the status above); the rerun uses the same harness as every other system.
Every system gets each question's chat history, then returns context that the same reader model uses to answer with the official LongMemEval prompt. Answers are graded by the official judge (gpt-4o-2024-08-06) and cross-checked by 3 other judges. Full details are on the methodology page.
What the vendor reports
| Score | Variant | Reader | Judge | Note | Source |
|---|---|---|---|---|---|
| 71.2% | S | GPT-4o | GPT-4o, official prompts | Zep (paper). | arxiv.org/abs/2501.13956 |
| 90.2% | S | GPT-5.4 | GPT-5.4 with chain-of-thought grading | Zep Cloud. | www.getzep.com/research |
Vendor-reported scores and ours are measured differently, so a gap does not by itself mean either number is wrong. Our pilot, and the rerun, use openai/gpt-6-luna as the reader that writes each answer, gpt-4o-2024-08-06 with the official LongMemEval judge prompt, LongMemEval-S (cleaned, 2025-09) and 30 questions, with Graphiti 0.30.2. Scores move with the reader model, the judge model and its prompt, the dataset version, the number of questions, and whether a hosted platform or the open-source package is tested.
Frequently asked questions
What is Graphiti?
The open-source temporal knowledge-graph engine behind Zep. Facts carry validity intervals, so old facts can be invalidated when they change. It is developed by Zep and released under the Apache-2.0 license.
Why is there no Graphiti score on MemVerdict yet?
Pilot result withdrawn: our first adapter added each session as one episode, while Graphiti's own server adds each message as its own episode. Being rerun the canonical way.
What LongMemEval score does Zep report for Graphiti?
Reported: 71.2% (LongMemEval-S; reader GPT-4o; judge GPT-4o, official prompts; Zep (paper)) and 90.2% (LongMemEval-S; reader GPT-5.4; judge GPT-5.4 with chain-of-thought grading; Zep Cloud). These were measured with setups different from ours, so they are not directly comparable.
When will Graphiti results be published?
When the rerun finishes, this page and every comparison involving Graphiti will show its accuracy, cost and latency from the same harness, reader and judge as the other systems.
Compare Graphiti with
- Cognee vs GraphitiGraphiti: pilot result withdrawn, rerun in progress
- Full context vs GraphitiGraphiti: pilot result withdrawn, rerun in progress
- Graphiti vs HindsightGraphiti: pilot result withdrawn, rerun in progress
- Graphiti vs LangMemGraphiti: pilot result withdrawn, rerun in progress
- Graphiti vs Mem0Graphiti: pilot result withdrawn, rerun in progress
- Graphiti vs Plain RAGGraphiti: pilot result withdrawn, rerun in progress