Full context vs Plain RAG: accuracy, cost and latency
Two no-memory baselines: retrieval over past sessions against putting the whole history in the prompt. Independent results on the same 30 questions of LongMemEval-S (cleaned, 2025-09), with the same reader model and the official judge. Pilot sample, so accuracy is shown with 95% intervals.
Full context
Baseline, not a memory system- Vendor
- Baseline
- Version
- -
- License
- -
- Result
- 83.3% (95% CI 66–93%)
Plain RAG
Baseline, not a memory system- Vendor
- Baseline
- Version
- -
- License
- -
- Result
- 86.7% (95% CI 70–95%)
Verdict
Generated from the pilot data, n=30 per system- AccuracyStatistically tied on accuracy at n=30: Plain RAG 86.7% (95% CI 70–95%) vs Full context 83.3% (95% CI 66–93%). The 95% intervals overlap, so the 3.3-point gap could be noise.
- CostPlain RAG costs 2.4× less per 1,000 questions: $5.45 vs $13.1 for Full context.
- LatencySimilar latency: median 4.8s for Full context vs 4.8s for Plain RAG per question (p90 8.0s vs 7.4s).
- IngestionThe Full context baseline needs no ingestion step; Plain RAG takes 1s to ingest one chat history (median).
- ContextThe reader sees a median of 500k characters of context per question with Full context and 127k with Plain RAG.
Key numbers side by side
Bold marks the better value. Accuracy rows are not bolded when the 95% intervals overlap.
| Metric | Full context | Plain RAG |
|---|---|---|
| Accuracy (official judge) | 83.3% | 86.7% |
| 95% interval | 66–93% | 70–95% |
| Panel accuracy | 86.7% | 86.7% |
| Cost per 1,000 questions | $13.1 | $5.45 |
| of which memory side | $0.00 | $2.08 |
| of which answering | $13.1 | $3.37 |
| Latency p50 | 4.8s | 4.8s |
| Latency p90 | 8.0s | 7.4s |
| Ingestion per history (median) | none | 1s |
| Context given to reader (median chars) | 500k | 127k |
Full context and Plain RAG among all tested systems
Full context and Plain RAG are highlighted; other systems are muted for context, baselines lightest. Each name links to the system's page.
Accuracy
Higher is betterShare of 30 questions answered correctly (official judge). Whiskers: 95% interval.
- Cognee90.0%74–97%
- Hindsight90.0%74–97%
- Plain RAG86.7%70–95%
- Mem083.3%66–93%
- Full context83.3%66–93%
- LangMem33.3%19–51%
Not shown: Graphiti (pilot result withdrawn, rerun in progress).
Cost per 1,000 questions
Lower is betterUSD billed (prices as of 2026-10-08). Solid: memory side (ingestion and retrieval). Light: answering.
Not shown: Graphiti (pilot result withdrawn, rerun in progress).
Latency per question
Lower is betterRetrieval plus answer, in seconds. Solid: median (p50). Light: up to p90.
- LangMem3.5sp90 5.3s
- Cognee4.0sp90 8.1s
- Mem04.3sp90 7.4s
- Full context4.8sp90 8.0s
- Plain RAG4.8sp90 7.4s
- Hindsight5.4sp90 8.1s
Not shown: Graphiti (pilot result withdrawn, rerun in progress).
Ingestion time
Lower is betterMedian time to load one question's chat history into the system.
Not shown: Graphiti (pilot result withdrawn, rerun in progress).
Accuracy by question type
Difference is Full context minus Plain RAG, in percentage points. Each type has only a few pilot questions (n), so one question can move a row by 13 points or more.
| Question type | n | Full context | Plain RAG | Difference |
|---|---|---|---|---|
| Single-session (user) | 4 | 100% | 100% | 0 |
| Single-session (assistant) | 3 | 100% | 100% | 0 |
| Preferences | 2 | 100% | 100% | 0 |
| Multi-session | 8 | 75% | 75% | 0 |
| Knowledge update | 4 | 100% | 100% | 0 |
| Temporal reasoning | 8 | 75% | 88% | −13 pts in favour of Plain RAG |
| Abstention | 1 | 0% | 0% | 0 |
What each one is
Full context
No memory system. The entire chat history (about 100k tokens) is placed in the reader's prompt.
How we ran it
Official LongMemEval long-context setting with the official answer prompt.
Plain RAG
No memory system. Each past session is embedded once; the 10 most similar sessions are given to the reader in date order.
How we ran it
text-embedding-3-small, cosine similarity, top 10 sessions.
Frequently asked questions
Is Full context better than Plain RAG?
Statistically tied on accuracy at n=30: Plain RAG 86.7% (95% CI 70–95%) vs Full context 83.3% (95% CI 66–93%). The 95% intervals overlap, so the 3.3-point gap could be noise. With the 3-judge panel majority instead of the official judge: Full context 86.7%, Plain RAG 86.7%. No question type separates them by more than one question (each question type has only 1 to 8 questions in this pilot). These are pilot numbers (30 questions each); the full 500-question run will narrow the intervals.
Which is cheaper, Full context or Plain RAG?
Plain RAG costs 2.4× less per 1,000 questions: $5.45 vs $13.1 for Full context. Full context: $13.1 per 1,000 questions, all of it answering (no memory-side cost). Plain RAG: $5.45 per 1,000 questions, of which $2.08 (38%) is memory-side (ingestion and retrieval) and $3.37 is answering. Prices as of 2026-10-08.
Which is faster, Full context or Plain RAG?
Similar latency: median 4.8s for Full context vs 4.8s for Plain RAG per question (p90 8.0s vs 7.4s). The Full context baseline needs no ingestion step; Plain RAG takes 1s to ingest one chat history (median). Latency is measured per question as retrieval plus answering; ingestion is the time to load one question's chat history.
How were Full context and Plain RAG tested?
Both were run on the same 30 questions of LongMemEval-S (cleaned, 2025-09) as every other system, by the same harness. For each question, the system ingests that question's chat history, then retrieves context that one reader model (openai/gpt-6-luna) uses to answer with the official LongMemEval prompt. Answers are graded by the official judge (gpt-4o-2024-08-06) and cross-checked by 3 other judges; all judges agreed on 95% of graded answers. Costs are what the providers billed (prices as of 2026-10-08). Full context: Official LongMemEval long-context setting with the official answer prompt. Plain RAG: text-embedding-3-small, cosine similarity, top 10 sessions.
Which should I use, Full context or Plain RAG?
Both are baselines, not memory systems. Statistically tied on accuracy at n=30: Plain RAG 86.7% (95% CI 70–95%) vs Full context 83.3% (95% CI 66–93%). The 95% intervals overlap, so the 3.3-point gap could be noise. Plain RAG costs 2.4× less per 1,000 questions: $5.45 vs $13.1 for Full context. The reader sees a median of 500k characters of context per question with Full context and 127k with Plain RAG. These are pilot numbers (30 questions each); the full 500-question run will narrow the intervals.
More comparisons
Other comparisons with Full context
- Cognee vs Full contextStatistically tied on accuracy at n=30; Full context 11× cheaper
- Full context vs GraphitiGraphiti: pilot result withdrawn, rerun in progress
- Full context vs HindsightStatistically tied on accuracy at n=30; Full context 13× cheaper
- Full context vs LangMemFull context more accurate and 4.7× cheaper
- Full context vs Mem0Tied on accuracy (83.3% each); Full context 7.1× cheaper
Other comparisons with Plain RAG
- Cognee vs Plain RAGStatistically tied on accuracy at n=30; Plain RAG 26× cheaper
- Graphiti vs Plain RAGGraphiti: pilot result withdrawn, rerun in progress
- Hindsight vs Plain RAGStatistically tied on accuracy at n=30; Plain RAG 31× cheaper
- LangMem vs Plain RAGPlain RAG more accurate and 11× cheaper
- Mem0 vs Plain RAGStatistically tied on accuracy at n=30; Plain RAG 17× cheaper