MemVerdict

Pilot results: 30 of 500 questions per system. Gaps of a few points are within noise; 95% intervals are shown. The full run is in progress.

Memory systems we have tested

Each system ran the same 30 questions of LongMemEval-S (cleaned, 2025-09), with the same reader model and the official judge (gpt-4o-2024-08-06). These are pilot numbers, so every accuracy comes with its 95% interval. Baselines are included as reference points: they are not memory systems. Compare any two.

Memory systems

Baselines

Simple approaches with no memory layer. If a memory system cannot beat them, its extra cost and moving parts are hard to justify on this benchmark.

Coming soon

  • EverOS

    EverMind · not run yet

  • MemOS

    MemTensor · not run yet

  • Letta

    Letta · not run yet

  • Supermemory

    Supermemory · not run yet