Cognee vs LangMem: accuracy, cost and latency
Independent results on the same 30 questions of LongMemEval-S (cleaned, 2025-09), with the same reader model and the official judge. Pilot sample, so accuracy is shown with 95% intervals.
Verdict
Generated from the pilot data, n=30 per system- AccuracyCognee is more accurate: 90.0% (95% CI 74–97%) vs 33.3% (95% CI 19–51%) for LangMem. The 95% intervals do not overlap, even at n=30.
- CostLangMem costs 2.3× less per 1,000 questions: $61.6 vs $141 for Cognee.
- LatencySimilar latency: median 4.0s for Cognee vs 3.5s for LangMem per question (p90 8.1s vs 5.3s).
- IngestionCognee ingests a chat history faster: 2.8 min vs 8.5 min for LangMem (median per question).
- ContextThe reader sees a median of 83k characters of context per question with Cognee and 23k with LangMem.
Key numbers side by side
Bold marks the better value. Accuracy rows are not bolded when the 95% intervals overlap.
| Metric | Cognee | LangMem |
|---|---|---|
| Accuracy (official judge) | 90.0% | 33.3% |
| 95% interval | 74–97% | 19–51% |
| Panel accuracy | 90.0% | 26.7% |
| Cost per 1,000 questions | $141 | $61.6 |
| of which memory side | $139 | $61.0 |
| of which answering | $2.23 | $0.61 |
| Latency p50 | 4.0s | 3.5s |
| Latency p90 | 8.1s | 5.3s |
| Ingestion per history (median) | 2.8 min | 8.5 min |
| Context given to reader (median chars) | 83k | 23k |
Cognee and LangMem among all tested systems
Cognee and LangMem are highlighted; other systems are muted for context, baselines lightest. Each name links to the system's page.
Accuracy
Higher is betterShare of 30 questions answered correctly (official judge). Whiskers: 95% interval.
- Cognee90.0%74–97%
- Hindsight90.0%74–97%
- Plain RAG86.7%70–95%
- Mem083.3%66–93%
- Full context83.3%66–93%
- LangMem33.3%19–51%
Not shown: Graphiti (pilot result withdrawn, rerun in progress).
Cost per 1,000 questions
Lower is betterUSD billed (prices as of 2026-10-08). Solid: memory side (ingestion and retrieval). Light: answering.
Not shown: Graphiti (pilot result withdrawn, rerun in progress).
Latency per question
Lower is betterRetrieval plus answer, in seconds. Solid: median (p50). Light: up to p90.
- LangMem3.5sp90 5.3s
- Cognee4.0sp90 8.1s
- Mem04.3sp90 7.4s
- Full context4.8sp90 8.0s
- Plain RAG4.8sp90 7.4s
- Hindsight5.4sp90 8.1s
Not shown: Graphiti (pilot result withdrawn, rerun in progress).
Ingestion time
Lower is betterMedian time to load one question's chat history into the system.
Not shown: Graphiti (pilot result withdrawn, rerun in progress).
Accuracy by question type
Difference is Cognee minus LangMem, in percentage points. Each type has only a few pilot questions (n), so one question can move a row by 13 points or more.
| Question type | n | Cognee | LangMem | Difference |
|---|---|---|---|---|
| Single-session (user) | 4 | 100% | 75% | +25 pts in favour of Cognee |
| Single-session (assistant) | 3 | 100% | 0% | +100 pts in favour of Cognee |
| Preferences | 2 | 100% | 50% | +50 pts in favour of Cognee |
| Multi-session | 8 | 88% | 13% | +75 pts in favour of Cognee |
| Knowledge update | 4 | 100% | 50% | +50 pts in favour of Cognee |
| Temporal reasoning | 8 | 75% | 38% | +37 pts in favour of Cognee |
| Abstention | 1 | 100% | 0% | +100 pts in favour of Cognee |
What each one is
Cognee
Builds a knowledge graph plus vector index from documents and conversations, then retrieves graph context for a query.
How we ran it
Python SDK in-process with embedded defaults. One text document per session with the session date at the top, then cognify(). Search: Cognee's default HYBRID_COMPLETION with only_context=True and default top_k.
LangMem
A memory manager that asks an LLM to extract, consolidate and update memories in a LangGraph store.
How we ran it
create_memory_store_manager with default instructions, one invoke() per session, LangGraph InMemoryStore with OpenAI embeddings, store.search() with the default limit (10).
What the vendors report
Cognee
| Score | Variant | Reader | Judge | Note | Source |
|---|---|---|---|---|---|
| 90.0% | S (cleaned, 2025-09), 30-question pilot | openai/gpt-6-luna | gpt-4o-2024-08-06, official prompt | MemVerdict measurement, Cognee 1.6.3 | This page |
| None found | - | - | - | No self-reported LongMemEval score found. | - |
LangMem
| Score | Variant | Reader | Judge | Note | Source |
|---|---|---|---|---|---|
| 33.3% | S (cleaned, 2025-09), 30-question pilot | openai/gpt-6-luna | gpt-4o-2024-08-06, official prompt | MemVerdict measurement, LangMem 0.0.30 | This page |
| None found | - | - | - | No self-reported LongMemEval score found. | - |
Vendor-reported scores and ours are measured differently, so a gap does not by itself mean either number is wrong. Our pilot uses openai/gpt-6-luna as the reader that writes each answer, gpt-4o-2024-08-06 with the official LongMemEval judge prompt, LongMemEval-S (cleaned, 2025-09) and 30 questions, with each system set up as described above. Scores move with the reader model, the judge model and its prompt, the dataset version, the number of questions, and whether a hosted platform or the open-source package is tested.
Frequently asked questions
Is Cognee better than LangMem?
Cognee is more accurate: 90.0% (95% CI 74–97%) vs 33.3% (95% CI 19–51%) for LangMem. The 95% intervals do not overlap, even at n=30. With the 3-judge panel majority instead of the official judge: Cognee 90.0%, LangMem 26.7%. The largest gap by question type is Multi-session: 7 of 8 vs 1 of 8 correct, in Cognee's favour; each question type has only 1 to 8 questions in this pilot. These are pilot numbers (30 questions each); the full 500-question run will narrow the intervals.
Which is cheaper, Cognee or LangMem?
LangMem costs 2.3× less per 1,000 questions: $61.6 vs $141 for Cognee. Cognee: $141 per 1,000 questions, of which $139 (98%) is memory-side (ingestion and retrieval) and $2.23 is answering. LangMem: $61.6 per 1,000 questions, of which $61.0 (99%) is memory-side (ingestion and retrieval) and $0.61 is answering. Prices as of 2026-10-08.
Which is faster, Cognee or LangMem?
Similar latency: median 4.0s for Cognee vs 3.5s for LangMem per question (p90 8.1s vs 5.3s). Cognee ingests a chat history faster: 2.8 min vs 8.5 min for LangMem (median per question). Latency is measured per question as retrieval plus answering; ingestion is the time to load one question's chat history.
How were Cognee and LangMem tested?
Both were run on the same 30 questions of LongMemEval-S (cleaned, 2025-09) as every other system, by the same harness. For each question, the system ingests that question's chat history, then retrieves context that one reader model (openai/gpt-6-luna) uses to answer with the official LongMemEval prompt. Answers are graded by the official judge (gpt-4o-2024-08-06) and cross-checked by 3 other judges; all judges agreed on 95% of graded answers. Costs are what the providers billed (prices as of 2026-10-08). Cognee: Python SDK in-process with embedded defaults. One text document per session with the session date at the top, then cognify(). Search: Cognee's default HYBRID_COMPLETION with only_context=True and default top_k. LangMem: create_memory_store_manager with default instructions, one invoke() per session, LangGraph InMemoryStore with OpenAI embeddings, store.search() with the default limit (10).
Which should I use, Cognee or LangMem?
If recall accuracy over long chat histories is what matters most, Cognee did clearly better in this pilot. LangMem is cheaper, though: LangMem costs 2.3× less per 1,000 questions: $61.6 vs $141 for Cognee. Cognee is Apache-2.0 and LangMem is MIT licensed. These are pilot numbers (30 questions each); the full 500-question run will narrow the intervals.
More comparisons
Other comparisons with Cognee
- Cognee vs Full contextStatistically tied on accuracy at n=30; Full context 11× cheaper
- Cognee vs GraphitiGraphiti: pilot result withdrawn, rerun in progress
- Cognee vs HindsightTied on accuracy (90.0% each); Cognee 1.2× cheaper
- Cognee vs Mem0Statistically tied on accuracy at n=30; Mem0 1.5× cheaper
- Cognee vs Plain RAGStatistically tied on accuracy at n=30; Plain RAG 26× cheaper
Other comparisons with LangMem
- Full context vs LangMemFull context more accurate and 4.7× cheaper
- Graphiti vs LangMemGraphiti: pilot result withdrawn, rerun in progress
- Hindsight vs LangMemHindsight more accurate; LangMem 2.8× cheaper
- LangMem vs Mem0Mem0 more accurate; LangMem 1.5× cheaper
- LangMem vs Plain RAGPlain RAG more accurate and 11× cheaper