|
|
|
|
|
by divitsheth
10 days ago
|
|
We’re working on evals now. We ran a small preliminary LoCoMo test with a 2k-token retrieval budget: Almanac: 55.7%
BM25: 51.8%
Supermemory: 47.6%
Mem0: 60.6% LoCoMo is not quite our use case. It tests conversational memory, while Almanac is more for coding agents trying to find their way around a codebase. These results show output quality at the same token budget, not token savings yet. We’re working on coding agent evals that should be more representative. |
|