as a tip - models will always find a way to cheat, you will probably need to impose some restrictions on what they do / are able to access in the sandbox environment
It depends on what you're measuring. I agree that model resourcefulness is useful, but if you're trying to simulate real user sessions, then Claude looking at upstream Git and fetching the answer directly is somewhat worthless.
In my case, I'm trying to measure how coding agents perform under realistic scenarios when implementing tasks, as a proxy for how agents perform when used by actual users for those same tasks, so it's important to ensure the agents are behaving realistically instead of "cheating" and looking up answers.
Happy to share resources! I've been pretty deep in the space :)
Why spend tokens working through a solution when you can simply look it up?
Thoughts?
Thank you for sharing this blog it's a good read! You are definitely plugged in to the benchmark space! :)