Agent Memory Leaderboard
Agent Memory Leaderboard (AML) is an open, unified and reproducible evaluation platform for long-term memory systems and memory-enabled agents. It standardizes Add/Search interfaces and holds the Answer, evaluation, orchestration, scoring and publication workflow constant so textual, coding and multimodal memory systems can be compared fairly.
At a Glance
- Universities and academic researchers
- Research institutes and independent research teams
- Open-source memory-system maintainers
- Commercial AI memory products and API providers
- +1 more
AI Tools by Agent Memory Leaderboard
(1)Agent Memory Leaderboard
AI Agent Memory Benchmark
Discussions
No discussions yet
Be the first to start a discussion about Agent Memory Leaderboard
Latest News
Cycle 2 Agent Memory Challenge scheduled to open with three tracks and RMB 150,000 prize pool
Official site confirmed the first cycle had closed and Cycle 2 was expected to open September 20
CSIG announced the second Agent Memory Challenge and its host/organizing institutions
MemoraX AI ranked first in AML's inaugural commercial ranking, bringing attention to the new benchmark
Products & Services
Public ranking and evaluation platform comparing textual, multimodal and coding-agent memory systems under versioned datasets, fixed answer/evaluation settings and capability-level metrics.
Recurring public evaluation challenge for researchers, open-source maintainers and commercial product teams. Participants host Add and Search APIs; AML runs the standardized Answer, Eval, orchestration, review and publication process.
Public GitHub repository containing per-benchmark evaluation contracts, public answer/scoring behavior and runtime configuration for transparency and methodological review; it is not the production leaderboard service and excludes protected benchmark data and participant artifacts.
Market Position
AML positions itself as a neutral measurement and comparison layer rather than as a memory database or agent-memory product. It differentiates through a common Add/Search contract, fixed downstream Answer/Eval conditions, private evaluation and review, version traceability, and separate comparable boards; the inaugural commercial board compared systems including MemoraX, MemOS, Mem0, Vectorize, SuperMemory, TencentDB and NetEase submissions.
Founding Story
AML was created to address the lack of directly comparable agent-memory results: memory systems had been evaluated on different datasets, with different answer models, retrieval settings, judges and aggregation rules. The organizers' initial vision was an open benchmark that would provide broad coverage, controlled comparison and capability-level diagnosis; the repository says AML launched on July 29, 2026, with researchers from more than 20 universities and research organizations.
Business Model
Revenue Model
The public challenge is free to enter. Participants pay their own API, database, bandwidth and compute costs, while AML covers the unified Answer, Eval and evaluation-orchestration costs. No subscription, API usage price or other commercial revenue model is described in the reviewed sources.
Pricing Tiers
No registration fee; participant teams provide and operate their own Add/Search service and infrastructure.
Target Markets
- Universities and academic researchers
- Research institutes and independent research teams
- Open-source memory-system maintainers
- Commercial AI memory products and API providers
- Enterprise systems and teams building memory-enabled agents
- Comparing long-term memory systems across research methods and commercial APIs
- Evaluating long conversations, cross-session history, personal preferences, temporal events and continuous narratives
- Testing coding agents' retrieval and reuse of prior debugging, architecture and development experience
- Testing memory over image-rich or other multimodal tasks
- Diagnosing memory-system strengths, failure modes, governance and privacy behavior
- Submitting reproducible open-source methods or verifiable hosted memory products