|
|
|
|
|
by docheinestages
31 days ago
|
|
We really need a "memory arena" to serve two important purposes: 1. List all the known agent memory projects (of which there are hundreds)\
2. Objectively compare and score them both against each other and vanilla harnesses like Claude Code Only then can I have the cognitive capacity to decide which one makes sense for me. |
|
Currently constructing a repeatable test harness for PMB: Fixed task, with/without memory, repeated N times, giving number of tokens/turns/passed/not passed with a subjective quality score too. Would be happy to share the task set and evaluation criteria for testing on anyone else's memory server or clean slate control, not just mine.