Hacker News new | ask | show | jobs
by Espressosaurus 1 day ago
I’ve observed the degradation, but I suspect what’s happening is they’re tuning it for lower inference costs. Maybe turning down the amount of thinking, maybe quantizing, maybe something else.

It seems like there’s a week by week and sometimes day by day change in performance when on a subscription plan using their harnesses.

2 comments

https://marginlab.ai/trackers/claude-code/ their tracker generally shows that isn’t the case. The only times I’ve seen it drop is something broken and just before fable launched.
Is this using the api or using a subscription, though? The incentives are different for each, and it isn't the least bit unexpected that they would maintain API access quality while 'optimizing' the subscription experience to improve their margins (or losses)

It seems to do really this you would need to crowdsource it -- users individually give the lab access to a body of subscriptions normally used by average people, and the lab occasionally runs some masked version of the task through on diverse accounts.

I mean they could just be routing known benchmark questions (which all of SWEBench are) to a full-performance variant.
i thought i had noticed a degradation, but it turned out claude code had swapped itself back to opus.

might be the case for you as well