Hacker News new | ask | show | jobs
by irthomasthomas 27 days ago
Why does anthropic change the set of benchmarks they use with every new model release?

https://www.anthropic.com/news/claude-opus-4-7

https://www.anthropic.com/news/claude-opus-4-6

1 comments

1. Benchmarks saturate 2. They select the most impressive improvments