|
|
|
|
|
by carterschonwald
106 days ago
|
|
omg this is so cool.
because im writing my own harness and i need some cognitive benchmarks. i have a bunch of harness level infra around llm interactions that seems to help with reasoning, but i dont have a structured way evaluate things thx for sharing your test setup, i really appreciate the time you took. this will help me so much |
|