Hindsight vs Plain RAG: accuracy, cost and latency
A memory system against a no-memory baseline: does the memory layer earn its cost? Independent results on the same 30 questions of LongMemEval-S (cleaned, 2025-09), with the same reader model and the official judge. Pilot sample, so accuracy is shown with 95% intervals.
Verdict
Generated from the pilot data, n=30 per systemIs a memory layer worth it here?
Not demonstrated in this pilot: Hindsight and the Plain RAG baseline are statistically tied on accuracy at n=30 (Hindsight 90.0% vs Plain RAG 86.7%; the 95% intervals overlap), and Hindsight costs 31× more per 1,000 questions. The reader sees a median of 18k characters of context per question with Hindsight and 127k with Plain RAG.
- AccuracyStatistically tied on accuracy at n=30: Hindsight 90.0% (95% CI 74–97%) vs Plain RAG 86.7% (95% CI 70–95%). The 95% intervals overlap, so the 3.3-point gap could be noise.
- CostPlain RAG costs 31× less per 1,000 questions: $5.45 vs $172 for Hindsight.
- LatencySimilar latency: median 5.4s for Hindsight vs 4.8s for Plain RAG per question (p90 8.1s vs 7.4s).
- IngestionPlain RAG ingests a chat history faster: 1s vs 8.8 min for Hindsight (median per question).
Key numbers side by side
Bold marks the better value. Accuracy rows are not bolded when the 95% intervals overlap.
| Metric | Hindsight | Plain RAG |
|---|---|---|
| Accuracy (official judge) | 90.0% | 86.7% |
| 95% interval | 74–97% | 70–95% |
| Panel accuracy | 86.7% | 86.7% |
| Cost per 1,000 questions | $172 | $5.45 |
| of which memory side | $171 | $2.08 |
| of which answering | $0.71 | $3.37 |
| Latency p50 | 5.4s | 4.8s |
| Latency p90 | 8.1s | 7.4s |
| Ingestion per history (median) | 8.8 min | 1s |
| Context given to reader (median chars) | 18k | 127k |
Hindsight and Plain RAG among all tested systems
Hindsight and Plain RAG are highlighted; other systems are muted for context, baselines lightest. Each name links to the system's page.
Accuracy
Higher is betterShare of 30 questions answered correctly (official judge). Whiskers: 95% interval.
- Cognee90.0%74–97%
- Hindsight90.0%74–97%
- Plain RAG86.7%70–95%
- Mem083.3%66–93%
- Full context83.3%66–93%
- LangMem33.3%19–51%
Not shown: Graphiti (pilot result withdrawn, rerun in progress).
Cost per 1,000 questions
Lower is betterUSD billed (prices as of 2026-10-08). Solid: memory side (ingestion and retrieval). Light: answering.
Not shown: Graphiti (pilot result withdrawn, rerun in progress).
Latency per question
Lower is betterRetrieval plus answer, in seconds. Solid: median (p50). Light: up to p90.
- LangMem3.5sp90 5.3s
- Cognee4.0sp90 8.1s
- Mem04.3sp90 7.4s
- Full context4.8sp90 8.0s
- Plain RAG4.8sp90 7.4s
- Hindsight5.4sp90 8.1s
Not shown: Graphiti (pilot result withdrawn, rerun in progress).
Ingestion time
Lower is betterMedian time to load one question's chat history into the system.
Not shown: Graphiti (pilot result withdrawn, rerun in progress).
Accuracy by question type
Difference is Hindsight minus Plain RAG, in percentage points. Each type has only a few pilot questions (n), so one question can move a row by 13 points or more.
| Question type | n | Hindsight | Plain RAG | Difference |
|---|---|---|---|---|
| Single-session (user) | 4 | 100% | 100% | 0 |
| Single-session (assistant) | 3 | 33% | 100% | −67 pts in favour of Plain RAG |
| Preferences | 2 | 100% | 100% | 0 |
| Multi-session | 8 | 88% | 75% | +13 pts in favour of Hindsight |
| Knowledge update | 4 | 100% | 100% | 0 |
| Temporal reasoning | 8 | 100% | 88% | +12 pts in favour of Hindsight |
| Abstention | 1 | 100% | 0% | +100 pts in favour of Hindsight |
What each one is
Hindsight
Extracts facts into a memory bank on Postgres, consolidates them into observations in the background, and recalls with hybrid search plus a local reranker.
How we ran it
Embedded daemon (pg0 Postgres), one bank per question. Sessions retained with their real timestamp; recall with the question date and default budget. We wait for background consolidation to finish before asking.
Plain RAG
No memory system. Each past session is embedded once; the 10 most similar sessions are given to the reader in date order.
How we ran it
text-embedding-3-small, cosine similarity, top 10 sessions.
What the vendors report
Hindsight
| Score | Variant | Reader | Judge | Note | Source |
|---|---|---|---|---|---|
| 90.0% | S (cleaned, 2025-09), 30-question pilot | openai/gpt-6-luna | gpt-4o-2024-08-06, official prompt | MemVerdict measurement, Hindsight 0.10.3 | This page |
| 91.4% | S | Gemini 3 Pro | GPT-OSS-120B | - | arxiv.org/abs/2512.12818 |
| 94.6% | S (cleaned) | Gemini 3.1 Pro | Gemini 2.5 Flash-Lite | - | benchmarks.hindsight.vectorize.io |
Vendor-reported scores and ours are measured differently, so a gap does not by itself mean either number is wrong. Our pilot uses openai/gpt-6-luna as the reader that writes each answer, gpt-4o-2024-08-06 with the official LongMemEval judge prompt, LongMemEval-S (cleaned, 2025-09) and 30 questions, with each system set up as described above. Scores move with the reader model, the judge model and its prompt, the dataset version, the number of questions, and whether a hosted platform or the open-source package is tested.
Frequently asked questions
Is Hindsight better than Plain RAG?
Statistically tied on accuracy at n=30: Hindsight 90.0% (95% CI 74–97%) vs Plain RAG 86.7% (95% CI 70–95%). The 95% intervals overlap, so the 3.3-point gap could be noise. With the 3-judge panel majority instead of the official judge: Hindsight 86.7%, Plain RAG 86.7%. The largest gap by question type is Single-session (assistant): 1 of 3 vs 3 of 3 correct, in Plain RAG's favour; each question type has only 1 to 8 questions in this pilot. These are pilot numbers (30 questions each); the full 500-question run will narrow the intervals.
Which is cheaper, Hindsight or Plain RAG?
Plain RAG costs 31× less per 1,000 questions: $5.45 vs $172 for Hindsight. Hindsight: $172 per 1,000 questions, of which $171 (over 99%) is memory-side (ingestion and retrieval) and $0.71 is answering. Plain RAG: $5.45 per 1,000 questions, of which $2.08 (38%) is memory-side (ingestion and retrieval) and $3.37 is answering. Prices as of 2026-10-08.
Which is faster, Hindsight or Plain RAG?
Similar latency: median 5.4s for Hindsight vs 4.8s for Plain RAG per question (p90 8.1s vs 7.4s). Plain RAG ingests a chat history faster: 1s vs 8.8 min for Hindsight (median per question). Latency is measured per question as retrieval plus answering; ingestion is the time to load one question's chat history.
How were Hindsight and Plain RAG tested?
Both were run on the same 30 questions of LongMemEval-S (cleaned, 2025-09) as every other system, by the same harness. For each question, the system ingests that question's chat history, then retrieves context that one reader model (openai/gpt-6-luna) uses to answer with the official LongMemEval prompt. Answers are graded by the official judge (gpt-4o-2024-08-06) and cross-checked by 3 other judges; all judges agreed on 95% of graded answers. Costs are what the providers billed (prices as of 2026-10-08). Hindsight: Embedded daemon (pg0 Postgres), one bank per question. Sessions retained with their real timestamp; recall with the question date and default budget. We wait for background consolidation to finish before asking. Plain RAG: text-embedding-3-small, cosine similarity, top 10 sessions.
Which should I use, Hindsight or Plain RAG?
Not demonstrated in this pilot: Hindsight and the Plain RAG baseline are statistically tied on accuracy at n=30 (Hindsight 90.0% vs Plain RAG 86.7%; the 95% intervals overlap), and Hindsight costs 31× more per 1,000 questions. The reader sees a median of 18k characters of context per question with Hindsight and 127k with Plain RAG. If the Plain RAG baseline is as accurate on your own data, it is the simpler option to run; test both on a sample of your real conversations before committing. These are pilot numbers (30 questions each); the full 500-question run will narrow the intervals.
More comparisons
Other comparisons with Hindsight
- Cognee vs HindsightTied on accuracy (90.0% each); Cognee 1.2× cheaper
- Full context vs HindsightStatistically tied on accuracy at n=30; Full context 13× cheaper
- Graphiti vs HindsightGraphiti: pilot result withdrawn, rerun in progress
- Hindsight vs LangMemHindsight more accurate; LangMem 2.8× cheaper
- Hindsight vs Mem0Statistically tied on accuracy at n=30; Mem0 1.8× cheaper
Other comparisons with Plain RAG
- Cognee vs Plain RAGStatistically tied on accuracy at n=30; Plain RAG 26× cheaper
- Full context vs Plain RAGStatistically tied on accuracy at n=30; Plain RAG 2.4× cheaper
- Graphiti vs Plain RAGGraphiti: pilot result withdrawn, rerun in progress
- LangMem vs Plain RAGPlain RAG more accurate and 11× cheaper
- Mem0 vs Plain RAGStatistically tied on accuracy at n=30; Plain RAG 17× cheaper