Hindsight: LongMemEval results
Extracts facts into a memory bank on Postgres, consolidates them into observations in the background, and recalls with hybrid search plus a local reranker.
- Vendor
- Vectorize
- Version tested
- 0.10.3
- License
- MIT
- Repository
- github.com/vectorize-io/hindsight
Headline numbers
Pilot: 30 questions of LongMemEval-S (cleaned, 2025-09). With this few questions, gaps of several points between systems are within noise, so read accuracy together with its 95% interval.
- Accuracy (official judge)
- 90.0%
- 95% CI 74–97% · 27/30 correct
- Rank by accuracy
- #1 of 6
- Among all scored systems, baselines included. Tied with Cognee. #1 of 4 memory systems.
- Panel accuracy (3-judge majority)
- 86.7%
- Cross-check with gpt-6-sol, claude-sonnet-5.5, gemini-3.1-pro-preview.
- Cost per 1,000 questions
- $172
- Memory side $171 · answering $0.71 · prices as of 2026-10-08
- Latency per question
- 5.4s p50
- p90 8.1s · retrieval plus answer
- Ingestion and context
- 8.8 min
- Median time to ingest one chat history; the reader sees a median 18k characters of context.
How Hindsight compares
All scored systems on the same questions, reader and judge. Hindsight is highlighted. Baselines are shown in a lighter grey.
Accuracy
Higher is betterShare of 30 questions answered correctly (official judge). Whiskers: 95% interval.
- Cognee90.0%74–97%
- Hindsight90.0%74–97%
- Plain RAG86.7%70–95%
- Mem083.3%66–93%
- Full context83.3%66–93%
- LangMem33.3%19–51%
Not shown: Graphiti (pilot result withdrawn, rerun in progress).
Cost per 1,000 questions
Lower is betterUSD billed (prices as of 2026-10-08). Solid: memory side (ingestion and retrieval). Light: answering.
Not shown: Graphiti (pilot result withdrawn, rerun in progress).
Latency per question
Lower is betterRetrieval plus answer, in seconds. Solid: median (p50). Light: up to p90.
- LangMem3.5sp90 5.3s
- Cognee4.0sp90 8.1s
- Mem04.3sp90 7.4s
- Full context4.8sp90 8.0s
- Plain RAG4.8sp90 7.4s
- Hindsight5.4sp90 8.1s
Not shown: Graphiti (pilot result withdrawn, rerun in progress).
Ingestion time
Lower is betterMedian time to load one question's chat history into the system.
Not shown: Graphiti (pilot result withdrawn, rerun in progress).
Accuracy by question type
LongMemEval groups questions by the memory ability they test. n is the number of pilot questions in each group. Highest: Abstention (n=1), Knowledge update (n=4), Preferences (n=2), Single-session (user) (n=4), Temporal reasoning (n=8) at 100%. Lowest: Single-session (assistant) (n=3) at 33%. Category samples are small, so treat these as hints.
- Single-session (user) n=4100% (4/4)
- Single-session (assistant) n=333% (1/3)
- Preferences n=2100% (2/2)
- Multi-session n=888% (7/8)
- Knowledge update n=4100% (4/4)
- Temporal reasoning n=8100% (8/8)
- Abstention n=1100% (1/1)
How we ran it
Embedded daemon (pg0 Postgres), one bank per question. Sessions retained with their real timestamp; recall with the question date and default budget. We wait for background consolidation to finish before asking.
Every system gets each question's chat history, then returns context that the same reader model (openai/gpt-6-luna) uses to answer with the official LongMemEval prompt. Answers are graded by the official judge (gpt-4o-2024-08-06) and cross-checked by 3 other judges. Full details are on the methodology page.
What the vendor reports
| Score | Variant | Reader | Judge | Note | Source |
|---|---|---|---|---|---|
| 90.0% | S (cleaned, 2025-09), 30-question pilot | openai/gpt-6-luna | gpt-4o-2024-08-06, official prompt | MemVerdict measurement, Hindsight 0.10.3 | This page |
| 91.4% | S | Gemini 3 Pro | GPT-OSS-120B | - | arxiv.org/abs/2512.12818 |
| 94.6% | S (cleaned) | Gemini 3.1 Pro | Gemini 2.5 Flash-Lite | - | benchmarks.hindsight.vectorize.io |
Vendor-reported scores and ours are measured differently, so a gap does not by itself mean either number is wrong. Our pilot uses openai/gpt-6-luna as the reader that writes each answer, gpt-4o-2024-08-06 with the official LongMemEval judge prompt, LongMemEval-S (cleaned, 2025-09) and 30 questions, with Hindsight 0.10.3 set up as described above. Scores move with the reader model, the judge model and its prompt, the dataset version, the number of questions, and whether a hosted platform or the open-source package is tested.
Frequently asked questions
How accurate is Hindsight on LongMemEval?
Hindsight answered 27 of 30 questions correctly (90.0%, 95% CI 74–97%) under the official judge. Rank 1 of 6 scored systems including baselines (tied); 1 of 4 memory systems (tied). The 3-judge panel majority gives 86.7%. This is a 30-question pilot of LongMemEval-S (cleaned, 2025-09); the full 500-question run is in progress.
How much does Hindsight cost to run?
Hindsight: $172 per 1,000 questions, of which $171 (over 99%) is memory-side (ingestion and retrieval) and $0.71 is answering. Memory-side cost covers ingesting each question's chat history and retrieving from it, as billed by the providers (prices as of 2026-10-08).
How fast is Hindsight?
Median latency is 5.4s per question (p90 8.1s), measured as retrieval plus answering. Ingesting one question's chat history takes 8.8 min (median).
Why is this Hindsight score different from the one Vectorize reports?
Vectorize reports 91.4% (LongMemEval-S; reader Gemini 3 Pro; judge GPT-OSS-120B) and 94.6% (LongMemEval-S (cleaned); reader Gemini 3.1 Pro; judge Gemini 2.5 Flash-Lite). We measured 90.0% (95% CI 74–97%). Vendor-reported scores and ours are measured differently, so a gap does not by itself mean either number is wrong. Our pilot uses openai/gpt-6-luna as the reader that writes each answer, gpt-4o-2024-08-06 with the official LongMemEval judge prompt, LongMemEval-S (cleaned, 2025-09) and 30 questions, with Hindsight 0.10.3 set up as described above. Scores move with the reader model, the judge model and its prompt, the dataset version, the number of questions, and whether a hosted platform or the open-source package is tested.
Does Hindsight beat the simple baselines?
Statistically tied on accuracy at n=30: Hindsight 90.0% (95% CI 74–97%) vs Plain RAG 86.7% (95% CI 70–95%). The 95% intervals overlap, so the 3.3-point gap could be noise. Statistically tied on accuracy at n=30: Hindsight 90.0% (95% CI 74–97%) vs Full context 83.3% (95% CI 66–93%). The 95% intervals overlap, so the 6.7-point gap could be noise.
Compare Hindsight with
- Cognee vs HindsightTied on accuracy (90.0% each); Cognee 1.2× cheaper
- Full context vs HindsightStatistically tied on accuracy at n=30; Full context 13× cheaper
- Graphiti vs HindsightGraphiti: pilot result withdrawn, rerun in progress
- Hindsight vs LangMemHindsight more accurate; LangMem 2.8× cheaper
- Hindsight vs Mem0Statistically tied on accuracy at n=30; Mem0 1.8× cheaper
- Hindsight vs Plain RAGStatistically tied on accuracy at n=30; Plain RAG 31× cheaper