MemVerdict

Pilot results: 30 of 500 questions per system. Gaps of a few points are within noise; 95% intervals are shown. The full run is in progress.

LongMemEval: the benchmark behind the scores

500 questions about long chat histories, each history around 101k tokens. Here is what the benchmark asks, how it was built, and what it can and cannot tell you about a memory system.

PaperICLR 2025, arXiv 2410.10813ReleaseCleaned, Sept 2025LicenseMIT
On this page

What LongMemEval is

LongMemEval is a public benchmark for the long-term memory of chat assistants, introduced by Di Wu et al. at ICLR 2025 (arXiv 2410.10813) and released under the MIT license on GitHub. It asks a simple question: after many conversations with a user, can an assistant still answer questions that depend on what was said?

Each of the 500 questions comes with its own chat history. A system ingests that history, then gets the question and must answer from what it remembers. The answer is graded against a reference answer by an LLM judge using the benchmark's official prompts. MemVerdict uses the cleaned release, September 2025.

How it is built

Every history is a sequence of dated chat sessions. Two kinds of session are mixed together:

  • Evidence sessions: simulated conversations in which the user mentions the facts the question depends on.
  • Filler sessions: conversations from the public ShareGPT and UltraChat datasets and sessions of other simulated users, which have nothing to do with the question.

In variant S, a history has about 48 sessions and 494 turns, around 101k tokens on average (o200k tokenizer; from 95k to 104k). Each question also has its own date, which matters for questions about time. The facts a question needs are a small part of a long, mostly unrelated history.

Question types

Questions fall into six types. Four of them also include abstention questions, where the correct answer is that the history does not say. Counts are for LongMemEval-S.

LongMemEval-S question types, counts and what each tests
TypeQuestionsAbstentionWhat it tests
Single-session (user)706Recall of a detail the user stated once, in one session.
Single-session (assistant)56–Recall of something the assistant itself said, not the user.
Preferences30–Using a preference the user expressed to personalize a new answer.
Multi-session13312Combining or counting facts spread over several sessions.
Temporal reasoning1336Reasoning about dates and order, using session dates and times mentioned in the chats.
Knowledge update786Answering with the latest value of a fact that changed over time.
Total50030

Illustrative examples

Single-session (user)
Weeks earlier the user mentioned adopting a beagle named Pico.“What is my dog's name?”
Single-session (assistant)
The assistant once recommended three science-fiction novels.“What was the second book you recommended to me?”
Preferences
The user said they are vegetarian and like spicy food.“Any idea what I could cook tonight?”
Multi-session
Concerts are mentioned in four different sessions.“How many concerts did I go to this spring?”
Temporal reasoning
The user started a pottery class in one session and mentioned a first exhibition later.“How many weeks after starting pottery did I have my first exhibition?”
Knowledge update
The user worked at a bakery, then said in a later session they had moved to a bookshop.“Where do I work now?”

Abstention

30 of the 500 questions ask about something the history never mentions. A system passes only if it says it does not know. For example (again illustrative), a user who only ever talked about a dog asks “What is my cat's name?”. Abstention checks that a memory system does not invent an answer from loosely related memories.

Variants S and M

LongMemEval variants S and M
VariantSessions per questionTokens per questionFits in reader contextOn MemVerdict
S~48~101kYesCurrent track
M~500~1.1MNo (reader: 1.05M tokens)Planned hard track

M has about ten times as many sessions per question as S. Its history is longer than our reader's 1.05M-token context window, so the reader cannot simply read everything.

Why S now fits in context

A 100k-token history now fits in the context window of today's long-context models, so a baseline that just puts the whole history in the prompt is strong on S. In the MemVerdict pilot, Full context scored 83.3% and Plain RAG 86.7%, judged by the official judge on 30 questions.

We report this openly rather than hide it. On S, cost, latency and the size of the context a system hands the reader matter alongside accuracy. Variant M is where memory becomes necessary.

Planned tracks

  • LongMemEval-M, the hard track: about 500 sessions (1.1M tokens) per question, longer than the reader's context.
  • MemoryAgentBench, conflict-resolution split, which focuses on facts that are later contradicted or updated.

Both are planned to run under the same methodology as the current track.

Citation and license

LongMemEval is the work of its authors; MemVerdict only runs it. If you use the benchmark, cite the paper:

Di Wu et al. “LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory.” ICLR 2025. arXiv:2410.10813.

The dataset and code are released under the MIT license at github.com/xiaowu0162/LongMemEval.