MemVerdict

Pilot results: 30 of 500 questions per system. Gaps of a few points are within noise; 95% intervals are shown. The full run is in progress.

How MemVerdict tests memory systems

Every system answers the same questions with the same reader, uses the same internal model and embeddings, is fed the way its own maintainers feed it, and is graded by the official LongMemEval judge. This page describes each of those choices and why it was made.

BenchmarkLongMemEval-S (cleaned, 2025-09)PhasePilot, 30 of 500 questionsPrices as of2026-10-08
On this page

Overview

MemVerdict measures how well long-term memory systems for AI agents work: whether an agent using a given memory system can answer questions about a long chat history, what that costs, and how long it takes. The test is LongMemEval, a public academic benchmark of 500 questions, each with its own history of about 100,000 tokens of chat.

Two baselines without any memory system run under the same rules: Full context puts the whole history in the reader's prompt, and Plain RAG embeds each past session and passes the 10 most similar ones to the reader. They show what a memory system adds over doing nothing special.

Settings at a glance

Benchmark settings at a glance
SettingValue
BenchmarkLongMemEval-S, cleaned release, September 2025, 500 questions
Current phasePilot: 30 questions per system, stratified by question type
Readeropenai/gpt-6-luna via OpenRouter, default settings, official LongMemEval answer prompt
Memory systems' internal modelopenai/gpt-6-luna, reasoning effort none, the same for every system
Embeddingstext-embedding-3-small for every system
Headline judgegpt-4o-2024-08-06 with the official per-type prompts, unchanged
Judge panelopenai/gpt-6-sol, anthropic/claude-sonnet-5.5, google/gemini-3.1-pro-preview; official prompts, reasoning effort low, majority vote
IngestionAs each system's own server, documentation or benchmark code does it; one isolated user per question
DatesProcess clock set to each session's date during ingestion and to the question date when asking
CostBilled cost incl. prompt-cache savings, prices as of 2026-10-08; list price kept as an upper bound
LatencyRetrieval + answer wall time per question, median and p90; Apple M5, 24 GB
Intervals95% Wilson score intervals
Timeout2 hours per question

Benchmark and question types

MemVerdict uses LongMemEval (Di Wu et al., ICLR 2025, MIT license) in its cleaned release, September 2025. Variant S has 500 questions. Each question comes with its own chat history of about 48 sessions and 494 turns, roughly 101k tokens. The histories mix sessions that contain the facts needed to answer with filler sessions from public chat datasets and other simulated users.

LongMemEval-S question types
Question typeQuestionsOf which abstention
Single-session (user)706
Single-session (assistant)56–
Preferences30–
Multi-session13312
Temporal reasoning1336
Knowledge update786
Total50030

Abstention questions ask about something the history never says; the right answer is to say so. What each type tests, with illustrative examples, is on the LongMemEval page.

The pilot sample

The pilot uses 30 questions per system, picked by deterministic proportional stratification over question types: the same questions for every system, in the same proportions as the full dataset.

Why S alone is not enough

A 100k-token history now fits in the context window of current long-context models, so simply reading the whole history is a strong baseline on S: in the pilot, Full context scored 83.3% and Plain RAG 86.7%. We report this openly rather than hide it. Variant M, with about 500 sessions (1.1M tokens) per question, is longer than the reader's context, so memory becomes necessary there. M is planned as the hard track.

What is held constant

A score should reflect the memory system, not the model behind it. Everything that is not the memory design is therefore the same for every system:

  • Reader. One model answers every question: openai/gpt-6-luna via OpenRouter with default settings. The answer prompt is LongMemEval's official one, imported verbatim from the official repository: the history template with step-by-step reading for the baselines that pass raw chat history, and the facts template for memory systems.
  • Internal model. Memory systems that call a language model to extract, consolidate or search memories all use openai/gpt-6-luna with reasoning effort fixed to none.
  • Embeddings. text-embedding-3-small for every system.
  • Local components at their defaults. Parts that run locally and are specific to one system keep their default configuration, for example Hindsight's local cross-encoder reranker and Mem0's BM25 and spaCy extras (as in Mem0's own benchmark).
  • Prompts. No system prompt is edited.

How each system is fed

Each system is fed the way its own server, documentation or benchmark code feeds it. If a system's server adds one message at a time, our adapter does too; if its benchmark code adds one user and assistant pair per call, so do we. The exact setup for each system is on its system page.

Each question is an isolated user. Every question gets its own memory instance and its own worker process, so nothing learned for one question can leak into another.

Clock simulation

LongMemEval sessions carry dates spread over months, and many questions depend on them. Many memory systems stamp what they store with the current time. To keep the dates right without changing any system's code:

  • while a session is being ingested, the process clock reads that session's date;
  • the question is asked with the clock set to the question's date;
  • systems that accept a timestamp parameter also receive the date natively.

Judging

Headline score: the official judge. Every answer is graded by LongMemEval's official judge, gpt-4o-2024-08-06, with the official per-question-type prompts, unchanged.

Check: a three-model panel. The same answers are also graded by openai/gpt-6-sol, anthropic/claude-sonnet-5.5, google/gemini-3.1-pro-preview, each with the same official prompts and reasoning effort set to low, and the panel decides by majority vote. In the pilot, all four judges agreed on 95% of answers.

Many published vendor numbers use a different judge or a more lenient custom prompt. That is one of the main reasons they differ from ours, and why each system page lists vendor-reported scores together with the reader and judge behind them.

Cost

Cost is the billed cost reported by the provider (OpenRouter's usage.cost, or OpenAI's usage), so it includes any savings a system gets from its own use of prompt caching. The list-price cost without cache discounts is kept as an upper bound. Prices are as of 2026-10-08.

Cost per 1,000 questions is

(memory ingestion + retrieval + answer) per question × 1,000

LongMemEval asks one question per history, so the cost of ingesting a history is never spread over many queries as it would be in real use. For that reason memory cost and answer cost are reported separately as well as combined.

Latency and ingestion time

Latency is the wall time to retrieve memories and answer one question, reported as the median and the 90th percentile. Ingestion time is the wall time to ingest one full history.

Everything is measured on one machine (Apple M5, 24 GB), with each system's embedded databases running locally and model calls going to OpenRouter over the internet. Absolute times depend on that setup and on network conditions; they are most useful for comparing systems with each other.

Statistics

Accuracy is shown with a 95% Wilson score interval, which stays well-behaved at small samples and near 0% or 100%. For k correct answers out of n, with p = k/n and z = 1.96:

(p + z²/2n ± z·√(p(1−p)/n + z²/4n²)) / (1 + z²/n)

At the pilot's n = 30, that interval is roughly ±13 points, so the pilot cannot rank systems whose scores are close: when two intervals overlap substantially, treat the systems as not yet distinguishable. With all 500 questions the intervals become about four times narrower.

Per-type scores in the pilot rest on a handful of questions each; read them as orientation, not as measurements.

Isolation and metering

  • Metering proxy. Every model call from every system, including the reader and the judges, goes through a proxy that records it and its billed cost.
  • No real keys. Systems never see a real API key.
  • No telemetry. Vendor telemetry is disabled for every system.
  • One process per question. Each question runs in its own worker process with its own memory instance.
  • Timeout. Each question has a limit of 2 hours.

Fairness and vendor review

The aim is to show each system as its maintainers intend it to be used, under rules that are the same for everyone. Before full results are published, the maintainers of each system are invited to review their configuration. They may propose better settings, which are then run under the same rules as everything else.

MemVerdict is independent: no memory vendor funds or runs it. See About for funding and conflicts of interest.

Corrections and changelog

Errors are corrected publicly. When a result is wrong because of our setup, it is withdrawn or fixed and the change is recorded here.

Pilot, October 2026 · Withdrawn

Graphiti pilot result withdrawn

Our adapter fed each session to Graphiti as one episode, while Graphiti's own server adds each message as an episode. The pilot result is withdrawn and Graphiti is being rerun the canonical way.

Reproducibility

With the full run, the harness, the adapter for each system, the raw answers, the retrieved contexts and every judge output will be released as open source, so anyone can check a single answer or rerun a system.

Submit a system

Any long-term memory system can be added. A system enters through an adapter that feeds it the way its own server, documentation or benchmark code does, and it runs under exactly the rules on this page: the same reader, internal model, embeddings, judges and isolation.

  • After the harness is released: adapters will be open source, so you can write or improve one and submit it for a run.
  • Until then: write to [email protected] with a link to your system and to the documentation or benchmark code that shows how it should be fed.

Already planned: EverOS, MemOS, Letta, Supermemory.