Hacker News new | ask | show | jobs
by xnorswap 1 day ago
And yet: https://marginlab.ai/trackers/claude-code-historical-perform...

There's clearly random variation, but it also shows each model release is just genuinely better. ( With the exception of a heavily degraded week of opus 4.7, which was acknowledged as a problem at the time. )

There's a psychology of getting used to models after being wowed by the new performance. It sets in as your new baseline expectations, and then when it doesn't deliver, it's felt more acutely. When it does deliver, it's just meeting expectations.

Then a new better model comes along and it's a step up again, another wow moment for a week or two until expectations adjust to meet the new baseline.

2 comments

I replied to the user above that referenced marginlab, but I believe marginlab uses the API. It is possible (arguably likely, in MBA-land) that the API and subscription accounts hit different sub-models.

Even if they use a subscription account, surely Anthropic can tell which one it is.

Remember https://en.wikipedia.org/wiki/Volkswagen_emissions_scandal? It's completely believable that benchmark-resembling requests are routes in a favorable manner.