Benchmark agent memory and RAG systems with MemoryBench 是做什麼的?
Use MemoryBench to run repeatable conversational memory and RAG benchmarks across providers, datasets, judge models, checkpoints, and structured reports.
Prerequisites
Bun, MemoryBench repository, at least one memory/RAG provider API key, at least one judge model API key, benchmark datasets
Installation
Basic usage or getting-started notes:
-
🆚 Multi‑provider comparison: run the same benchmark across providers side‑by‑side
-
📊 Structured reports: export run status, failures, and metrics for analysis
-
bun install
-
Extracted from upstream docs: https://raw.githubusercontent.com/supermemoryai/memorybench/HEAD/README.md