Full context vs Mem0: accuracy, cost and latency
A memory system against a no-memory baseline: does the memory layer earn its cost? Independent results on the same 30 questions of LongMemEval-S (cleaned, 2025-09), with the same reader model and the official judge. Pilot sample, so accuracy is shown with 95% intervals.
Full context
Baseline, not a memory system- Vendor
- Baseline
- Version
- -
- License
- -
- Result
- 83.3% (95% CI 66–93%)
Mem0
Memory system- Vendor
- Mem0
- Version
- 2.2.1 (open source)
- License
- Apache-2.0
- Result
- 83.3% (95% CI 66–93%)
Verdict
Generated from the pilot data, n=30 per systemIs a memory layer worth it here?
Not demonstrated in this pilot: Mem0 and the Full context baseline are statistically tied on accuracy at n=30 (Mem0 83.3% vs Full context 83.3%; the 95% intervals overlap), and Mem0 costs 7.1× more per 1,000 questions. The reader sees a median of 50k characters of context per question with Mem0 and 500k with Full context.
- AccuracyFull context and Mem0 are tied on accuracy: both answered 25 of 30 questions correctly (83.3%, 95% CI 66–93%).
- CostFull context costs 7.1× less per 1,000 questions: $13.1 vs $93.0 for Mem0.
- LatencySimilar latency: median 4.8s for Full context vs 4.3s for Mem0 per question (p90 8.0s vs 7.4s).
- IngestionThe Full context baseline needs no ingestion step; Mem0 takes 15.2 min to ingest one chat history (median).
Key numbers side by side
Bold marks the better value. Accuracy rows are not bolded when the 95% intervals overlap.
| Metric | Full context | Mem0 |
|---|---|---|
| Accuracy (official judge) | 83.3% | 83.3% |
| 95% interval | 66–93% | 66–93% |
| Panel accuracy | 86.7% | 83.3% |
| Cost per 1,000 questions | $13.1 | $93.0 |
| of which memory side | $0.00 | $91.6 |
| of which answering | $13.1 | $1.32 |
| Latency p50 | 4.8s | 4.3s |
| Latency p90 | 8.0s | 7.4s |
| Ingestion per history (median) | none | 15.2 min |
| Context given to reader (median chars) | 500k | 50k |
Full context and Mem0 among all tested systems
Full context and Mem0 are highlighted; other systems are muted for context, baselines lightest. Each name links to the system's page.
Accuracy
Higher is betterShare of 30 questions answered correctly (official judge). Whiskers: 95% interval.
- Cognee90.0%74–97%
- Hindsight90.0%74–97%
- Plain RAG86.7%70–95%
- Mem083.3%66–93%
- Full context83.3%66–93%
- LangMem33.3%19–51%
Not shown: Graphiti (pilot result withdrawn, rerun in progress).
Cost per 1,000 questions
Lower is betterUSD billed (prices as of 2026-10-08). Solid: memory side (ingestion and retrieval). Light: answering.
Not shown: Graphiti (pilot result withdrawn, rerun in progress).
Latency per question
Lower is betterRetrieval plus answer, in seconds. Solid: median (p50). Light: up to p90.
- LangMem3.5sp90 5.3s
- Cognee4.0sp90 8.1s
- Mem04.3sp90 7.4s
- Full context4.8sp90 8.0s
- Plain RAG4.8sp90 7.4s
- Hindsight5.4sp90 8.1s
Not shown: Graphiti (pilot result withdrawn, rerun in progress).
Ingestion time
Lower is betterMedian time to load one question's chat history into the system.
Not shown: Graphiti (pilot result withdrawn, rerun in progress).
Accuracy by question type
Difference is Full context minus Mem0, in percentage points. Each type has only a few pilot questions (n), so one question can move a row by 13 points or more.
| Question type | n | Full context | Mem0 | Difference |
|---|---|---|---|---|
| Single-session (user) | 4 | 100% | 75% | +25 pts in favour of Full context |
| Single-session (assistant) | 3 | 100% | 100% | 0 |
| Preferences | 2 | 100% | 100% | 0 |
| Multi-session | 8 | 75% | 75% | 0 |
| Knowledge update | 4 | 100% | 100% | 0 |
| Temporal reasoning | 8 | 75% | 88% | −13 pts in favour of Mem0 |
| Abstention | 1 | 0% | 0% | 0 |
What each one is
Full context
No memory system. The entire chat history (about 100k tokens) is placed in the reader's prompt.
How we ran it
Official LongMemEval long-context setting with the official answer prompt.
Mem0
Extracts short memories from each exchange with an LLM, stores them in a vector store, and retrieves them with semantic, keyword and entity signals.
How we ran it
Open-source SDK with local Qdrant, NLP extras installed (spaCy, BM25). Mirrors Mem0's own LongMemEval harness: one add() per user+assistant pair, top_k=200 at search. Historical timestamps are platform-only, so the shared clock simulation supplies dates.
What the vendors report
Mem0
| Score | Variant | Reader | Judge | Note | Source |
|---|---|---|---|---|---|
| 83.3% | S (cleaned, 2025-09), 30-question pilot | openai/gpt-6-luna | gpt-4o-2024-08-06, official prompt | MemVerdict measurement, Mem0 2.2.1 (open source) | This page |
| 94.4% | S | GPT-5 | GPT-5 with Mem0's own lenient prompt | Managed platform, not the open-source SDK. | mem0.ai/research |
Vendor-reported scores and ours are measured differently, so a gap does not by itself mean either number is wrong. Our pilot uses openai/gpt-6-luna as the reader that writes each answer, gpt-4o-2024-08-06 with the official LongMemEval judge prompt, LongMemEval-S (cleaned, 2025-09) and 30 questions, with each system set up as described above. Scores move with the reader model, the judge model and its prompt, the dataset version, the number of questions, and whether a hosted platform or the open-source package is tested.
Frequently asked questions
Is Full context better than Mem0?
Full context and Mem0 are tied on accuracy: both answered 25 of 30 questions correctly (83.3%, 95% CI 66–93%). With the 3-judge panel majority instead of the official judge: Full context 86.7%, Mem0 83.3%. No question type separates them by more than one question (each question type has only 1 to 8 questions in this pilot). These are pilot numbers (30 questions each); the full 500-question run will narrow the intervals.
Which is cheaper, Full context or Mem0?
Full context costs 7.1× less per 1,000 questions: $13.1 vs $93.0 for Mem0. Full context: $13.1 per 1,000 questions, all of it answering (no memory-side cost). Mem0: $93.0 per 1,000 questions, of which $91.6 (99%) is memory-side (ingestion and retrieval) and $1.32 is answering. Prices as of 2026-10-08.
Which is faster, Full context or Mem0?
Similar latency: median 4.8s for Full context vs 4.3s for Mem0 per question (p90 8.0s vs 7.4s). The Full context baseline needs no ingestion step; Mem0 takes 15.2 min to ingest one chat history (median). Latency is measured per question as retrieval plus answering; ingestion is the time to load one question's chat history.
How were Full context and Mem0 tested?
Both were run on the same 30 questions of LongMemEval-S (cleaned, 2025-09) as every other system, by the same harness. For each question, the system ingests that question's chat history, then retrieves context that one reader model (openai/gpt-6-luna) uses to answer with the official LongMemEval prompt. Answers are graded by the official judge (gpt-4o-2024-08-06) and cross-checked by 3 other judges; all judges agreed on 95% of graded answers. Costs are what the providers billed (prices as of 2026-10-08). Full context: Official LongMemEval long-context setting with the official answer prompt. Mem0: Open-source SDK with local Qdrant, NLP extras installed (spaCy, BM25). Mirrors Mem0's own LongMemEval harness: one add() per user+assistant pair, top_k=200 at search. Historical timestamps are platform-only, so the shared clock simulation supplies dates.
Which should I use, Full context or Mem0?
Not demonstrated in this pilot: Mem0 and the Full context baseline are statistically tied on accuracy at n=30 (Mem0 83.3% vs Full context 83.3%; the 95% intervals overlap), and Mem0 costs 7.1× more per 1,000 questions. The reader sees a median of 50k characters of context per question with Mem0 and 500k with Full context. If the Full context baseline is as accurate on your own data, it is the simpler option to run; test both on a sample of your real conversations before committing. These are pilot numbers (30 questions each); the full 500-question run will narrow the intervals.
More comparisons
Other comparisons with Full context
- Cognee vs Full contextStatistically tied on accuracy at n=30; Full context 11× cheaper
- Full context vs GraphitiGraphiti: pilot result withdrawn, rerun in progress
- Full context vs HindsightStatistically tied on accuracy at n=30; Full context 13× cheaper
- Full context vs LangMemFull context more accurate and 4.7× cheaper
- Full context vs Plain RAGStatistically tied on accuracy at n=30; Plain RAG 2.4× cheaper
Other comparisons with Mem0
- Cognee vs Mem0Statistically tied on accuracy at n=30; Mem0 1.5× cheaper
- Graphiti vs Mem0Graphiti: pilot result withdrawn, rerun in progress
- Hindsight vs Mem0Statistically tied on accuracy at n=30; Mem0 1.8× cheaper
- LangMem vs Mem0Mem0 more accurate; LangMem 1.5× cheaper
- Mem0 vs Plain RAGStatistically tied on accuracy at n=30; Plain RAG 17× cheaper