Plain RAG baseline: LongMemEval results
No memory system. Each past session is embedded once; the 10 most similar sessions are given to the reader in date order.
- Vendor
- Baseline
- Version tested
- -
- License
- -
Headline numbers
Pilot: 30 questions of LongMemEval-S (cleaned, 2025-09). With this few questions, gaps of several points between systems are within noise, so read accuracy together with its 95% interval.
- Accuracy (official judge)
- 86.7%
- 95% CI 70–95% · 26/30 correct
- Rank by accuracy
- #3 of 6
- Among all scored systems, baselines included.
- Panel accuracy (3-judge majority)
- 86.7%
- Cross-check with gpt-6-sol, claude-sonnet-5.5, gemini-3.1-pro-preview.
- Cost per 1,000 questions
- $5.45
- Memory side $2.08 · answering $3.37 · prices as of 2026-10-08
- Latency per question
- 4.8s p50
- p90 7.4s · retrieval plus answer
- Ingestion and context
- 1s
- Median time to ingest one chat history; the reader sees a median 127k characters of context.
How Plain RAG compares
All scored systems on the same questions, reader and judge. Plain RAG is highlighted. Baselines are shown in a lighter grey.
Accuracy
Higher is betterShare of 30 questions answered correctly (official judge). Whiskers: 95% interval.
- Cognee90.0%74–97%
- Hindsight90.0%74–97%
- Plain RAG86.7%70–95%
- Mem083.3%66–93%
- Full context83.3%66–93%
- LangMem33.3%19–51%
Not shown: Graphiti (pilot result withdrawn, rerun in progress).
Cost per 1,000 questions
Lower is betterUSD billed (prices as of 2026-10-08). Solid: memory side (ingestion and retrieval). Light: answering.
Not shown: Graphiti (pilot result withdrawn, rerun in progress).
Latency per question
Lower is betterRetrieval plus answer, in seconds. Solid: median (p50). Light: up to p90.
- LangMem3.5sp90 5.3s
- Cognee4.0sp90 8.1s
- Mem04.3sp90 7.4s
- Full context4.8sp90 8.0s
- Plain RAG4.8sp90 7.4s
- Hindsight5.4sp90 8.1s
Not shown: Graphiti (pilot result withdrawn, rerun in progress).
Ingestion time
Lower is betterMedian time to load one question's chat history into the system.
Not shown: Graphiti (pilot result withdrawn, rerun in progress).
Accuracy by question type
LongMemEval groups questions by the memory ability they test. n is the number of pilot questions in each group. Highest: Knowledge update (n=4), Single-session (assistant) (n=3), Preferences (n=2), Single-session (user) (n=4) at 100%. Lowest: Abstention (n=1) at 0%. Category samples are small, so treat these as hints.
- Single-session (user) n=4100% (4/4)
- Single-session (assistant) n=3100% (3/3)
- Preferences n=2100% (2/2)
- Multi-session n=875% (6/8)
- Knowledge update n=4100% (4/4)
- Temporal reasoning n=888% (7/8)
- Abstention n=10% (0/1)
How we ran it
text-embedding-3-small, cosine similarity, top 10 sessions.
Every system gets each question's chat history, then returns context that the same reader model (openai/gpt-6-luna) uses to answer with the official LongMemEval prompt. Answers are graded by the official judge (gpt-4o-2024-08-06) and cross-checked by 3 other judges. Full details are on the methodology page.
Frequently asked questions
What is the Plain RAG baseline?
No memory system. Each past session is embedded once; the 10 most similar sessions are given to the reader in date order. Setup: text-embedding-3-small, cosine similarity, top 10 sessions. We include it as a reference point for the memory systems.
How accurate is Plain RAG on LongMemEval?
Plain RAG answered 26 of 30 questions correctly (86.7%, 95% CI 70–95%) under the official judge. Rank 3 of 6 scored systems including baselines. The 3-judge panel majority gives 86.7%. This is a 30-question pilot of LongMemEval-S (cleaned, 2025-09); the full 500-question run is in progress.
How much does Plain RAG cost to run?
Plain RAG: $5.45 per 1,000 questions, of which $2.08 (38%) is memory-side (ingestion and retrieval) and $3.37 is answering. Memory-side cost covers ingesting each question's chat history and retrieving from it, as billed by the providers (prices as of 2026-10-08).
How fast is Plain RAG?
Median latency is 4.8s per question (p90 7.4s), measured as retrieval plus answering. Ingesting one question's chat history takes 1s (median).
Why include a Plain RAG baseline?
Baselines show whether a memory layer adds anything over a simple approach on the same questions, reader and judge. In this pilot, no memory system is clearly more accurate than Plain RAG, while Plain RAG is clearly more accurate than LangMem; the rest are statistically tied with it at n=30.
Compare Plain RAG with
- Cognee vs Plain RAGStatistically tied on accuracy at n=30; Plain RAG 26× cheaper
- Full context vs Plain RAGStatistically tied on accuracy at n=30; Plain RAG 2.4× cheaper
- Graphiti vs Plain RAGGraphiti: pilot result withdrawn, rerun in progress
- Hindsight vs Plain RAGStatistically tied on accuracy at n=30; Plain RAG 31× cheaper
- LangMem vs Plain RAGPlain RAG more accurate and 11× cheaper
- Mem0 vs Plain RAGStatistically tied on accuracy at n=30; Plain RAG 17× cheaper