What LongMemEval is
LongMemEval is a public benchmark for the long-term memory of chat assistants, introduced by Di Wu et al. at ICLR 2025 (arXiv 2410.10813) and released under the MIT license on GitHub. It asks a simple question: after many conversations with a user, can an assistant still answer questions that depend on what was said?
Each of the 500 questions comes with its own chat history. A system ingests that history, then gets the question and must answer from what it remembers. The answer is graded against a reference answer by an LLM judge using the benchmark's official prompts. MemVerdict uses the cleaned release, September 2025.
How it is built
Every history is a sequence of dated chat sessions. Two kinds of session are mixed together:
- Evidence sessions: simulated conversations in which the user mentions the facts the question depends on.
- Filler sessions: conversations from the public ShareGPT and UltraChat datasets and sessions of other simulated users, which have nothing to do with the question.
In variant S, a history has about 48 sessions and 494 turns, around 101k tokens on average (o200k tokenizer; from 95k to 104k). Each question also has its own date, which matters for questions about time. The facts a question needs are a small part of a long, mostly unrelated history.
Question types
Questions fall into six types. Four of them also include abstention questions, where the correct answer is that the history does not say. Counts are for LongMemEval-S.
| Type | Questions | Abstention | What it tests |
|---|---|---|---|
| Single-session (user) | 70 | 6 | Recall of a detail the user stated once, in one session. |
| Single-session (assistant) | 56 | – | Recall of something the assistant itself said, not the user. |
| Preferences | 30 | – | Using a preference the user expressed to personalize a new answer. |
| Multi-session | 133 | 12 | Combining or counting facts spread over several sessions. |
| Temporal reasoning | 133 | 6 | Reasoning about dates and order, using session dates and times mentioned in the chats. |
| Knowledge update | 78 | 6 | Answering with the latest value of a fact that changed over time. |
| Total | 500 | 30 |
Illustrative examples
- Single-session (user)
- Weeks earlier the user mentioned adopting a beagle named Pico.“What is my dog's name?”
- Single-session (assistant)
- The assistant once recommended three science-fiction novels.“What was the second book you recommended to me?”
- Preferences
- The user said they are vegetarian and like spicy food.“Any idea what I could cook tonight?”
- Multi-session
- Concerts are mentioned in four different sessions.“How many concerts did I go to this spring?”
- Temporal reasoning
- The user started a pottery class in one session and mentioned a first exhibition later.“How many weeks after starting pottery did I have my first exhibition?”
- Knowledge update
- The user worked at a bakery, then said in a later session they had moved to a bookshop.“Where do I work now?”
Abstention
30 of the 500 questions ask about something the history never mentions. A system passes only if it says it does not know. For example (again illustrative), a user who only ever talked about a dog asks “What is my cat's name?”. Abstention checks that a memory system does not invent an answer from loosely related memories.
Variants S and M
| Variant | Sessions per question | Tokens per question | Fits in reader context | On MemVerdict |
|---|---|---|---|---|
| S | ~48 | ~101k | Yes | Current track |
| M | ~500 | ~1.1M | No (reader: 1.05M tokens) | Planned hard track |
M has about ten times as many sessions per question as S. Its history is longer than our reader's 1.05M-token context window, so the reader cannot simply read everything.
Why S now fits in context
A 100k-token history now fits in the context window of today's long-context models, so a baseline that just puts the whole history in the prompt is strong on S. In the MemVerdict pilot, Full context scored 83.3% and Plain RAG 86.7%, judged by the official judge on 30 questions.
We report this openly rather than hide it. On S, cost, latency and the size of the context a system hands the reader matter alongside accuracy. Variant M is where memory becomes necessary.
Planned tracks
- LongMemEval-M, the hard track: about 500 sessions (1.1M tokens) per question, longer than the reader's context.
- MemoryAgentBench, conflict-resolution split, which focuses on facts that are later contradicted or updated.
Both are planned to run under the same methodology as the current track.
Citation and license
LongMemEval is the work of its authors; MemVerdict only runs it. If you use the benchmark, cite the paper:
Di Wu et al. “LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory.” ICLR 2025. arXiv:2410.10813.
The dataset and code are released under the MIT license at github.com/xiaowu0162/LongMemEval.