Hacker News new | ask | show | jobs
by Retric 46 days ago
I suspected you felt that way even though it hasn’t been my personal experience.

I’ve heard people say older models can’t do X, when I used that way etc. I suspect people are applying their own learning curve as part of their assessment of progress, you get better at writing prompts and it feels like the model improved.

Which is why I’m saying we need some objective metrics to judge predictions of actual capacity.

1 comments

There are objective metrics, that’s what benchmarks are.
Companies game benchmarks, but sure they are better than nothing.
I mean you wanted something objective, and they are. I don’t know why you’re being dismissive of them, they’re a huge element of what drives model development forward.

These companies aren’t just making stuff up, they really do want to improve the models, and the models really are improving.

> I don’t know why you’re being dismissive of them

I’m aware of multiple cases of benchmark cheating/“optimization”. So, taking benchmarks a face value seems laughable.