MemVerdict

Pilot results: 30 of 500 questions per system. Gaps of a few points are within noise; 95% intervals are shown. The full run is in progress.

About MemVerdict

Independent benchmark of long-term memory for AI agents. It is self-funded, no memory vendor funds or runs it, and every result is produced under the same published methodology.

On this page

Who runs MemVerdict

MemVerdict is an independent, self-funded project. Its purpose is to give people choosing a memory system for an AI agent numbers they can compare: the same benchmark, the same models and the same judge for every system, with cost and latency next to accuracy.

Independence

  • No memory vendor funds or runs MemVerdict.
  • Vendors are invited to review their system's configuration before full results are published and may propose better settings. Those settings are run under the same rules as every other system; see Fairness and vendor review.
  • Mistakes are corrected publicly in the changelog.

Funding

MemVerdict is self-funded. There are no sponsorships, no advertising and no affiliate links.

Conflicts of interest

The maintainers do not sell or build a memory system, and no maintainer works for a memory vendor.

If a maintainer ever builds a memory system, it will be clearly labeled as such on every page where it appears, evaluated with the same harness and rules as every other system, and disclosed on this page.

Using the results

MemVerdict results data is licensed CC BY 4.0: you may reuse it with attribution. The LongMemEval benchmark belongs to its authors and is released under the MIT license; see the LongMemEval page for the citation. The harness, adapters, raw answers, retrieved contexts and judge outputs will be released as open source with the full run.

Contact

For questions, corrections or to propose a system, write to [email protected]. Maintainers who want their system added can also read Submit a system.