Hacker News new | ask | show | jobs
by fxwin 16 days ago
I would expect they have production based datasets they evaluate new models against.
1 comments

I would not expect that. They wouldnt have missed mentioning it if that was the case. Its mostly driven by vibes.
>For four months, no frontier model beat Claude Opus in our production evals

What do you think that is referring to?