Hacker News new | ask | show | jobs
by sothatsit 4 days ago
This matches my experience of Opus 5 being a nice improvement over Opus 4.8, but not being revolutionary like Fable felt.

I’ve now replaced my use of Opus 4.8 xhigh with Opus 5 medium, and I’m using less tokens and it’s quicker. I can understand people being annoyed by its writing style but for getting work done that really doesn’t bother me. I’ve been really enjoying using it.

3 comments

I think they neutered Fable. When it first came out it was indeed revolutionary. But what we have today, is not what we had before the ban.
Noticed that too. I wonder if these things just degrade over time, perhaps with the way it writes memories about my project as it goes
I’ve observed the degradation, but I suspect what’s happening is they’re tuning it for lower inference costs. Maybe turning down the amount of thinking, maybe quantizing, maybe something else.

It seems like there’s a week by week and sometimes day by day change in performance when on a subscription plan using their harnesses.

https://marginlab.ai/trackers/claude-code/ their tracker generally shows that isn’t the case. The only times I’ve seen it drop is something broken and just before fable launched.
Is this using the api or using a subscription, though? The incentives are different for each, and it isn't the least bit unexpected that they would maintain API access quality while 'optimizing' the subscription experience to improve their margins (or losses)

It seems to do really this you would need to crowdsource it -- users individually give the lab access to a body of subscriptions normally used by average people, and the lab occasionally runs some masked version of the task through on diverse accounts.

I mean they could just be routing known benchmark questions (which all of SWEBench are) to a full-performance variant.
i thought i had noticed a degradation, but it turned out claude code had swapped itself back to opus.

might be the case for you as well

I feel like a conspiracy theorist but it feels like every new model release from both Anthropic and OpenAI has 1-2 weeks of fantastic performance then a gradual (and sometimes not so gradual) decline in intelligence. It's like they're quantising the model in the background to optimise for available compute/RAM. Which, if I put my MBA hat on, would make perfect sense. Why not halve the required RAM for "90%" of the performance? They get fantastic benchmarks and glowing reviews on release, then slowly squeeze more performance out of the model. By the time the next model is ready for release, the jump feels quite large again.
This is also my experience. I don't know if it's because of the quantization theory, or if it's just me getting used to a certain level of coding performance and gradually less tolerant of the mistakes it makes more over time.
And yet: https://marginlab.ai/trackers/claude-code-historical-perform...

There's clearly random variation, but it also shows each model release is just genuinely better. ( With the exception of a heavily degraded week of opus 4.7, which was acknowledged as a problem at the time. )

There's a psychology of getting used to models after being wowed by the new performance. It sets in as your new baseline expectations, and then when it doesn't deliver, it's felt more acutely. When it does deliver, it's just meeting expectations.

Then a new better model comes along and it's a step up again, another wow moment for a week or two until expectations adjust to meet the new baseline.

Remember https://en.wikipedia.org/wiki/Volkswagen_emissions_scandal? It's completely believable that benchmark-resembling requests are routes in a favorable manner.
I replied to the user above that referenced marginlab, but I believe marginlab uses the API. It is possible (arguably likely, in MBA-land) that the API and subscription accounts hit different sub-models.

Even if they use a subscription account, surely Anthropic can tell which one it is.

You should. Feel like a conspiracy theorist when saying things like this.

Users are not reliable or consistent model evaluators. Users adapt to models - the moment they get a model that performs better, their expectations rise, their tasks get harder and their prompts get shorter.

"They made the model worse" is PEBKAC in 9 cases out of 10.

There are few conspiracies where the vectors between "capitalist organization makes more money" and "user can't reliably distinguish tiers of product quality" overlap.
I wish the users weren't so fucking stupid with the "they made the model worse" stuff.

Then that 1 out of 10 case where the model was actually made worse (whether intentionally or by mistake) would stand out instead of being swallowed by the noise floor.

At this point I don't even bother with it. Constantly falls back to Opus anyway, so I may as well save myself some time.
Can you elaborate on what felt revolutionary to you about Fable?
I got Fable to run overnight and I woke up to a working prototype of a very complex feature. And then I did it again for another complex feature the next night.

The code still took weeks to clean up, but it worked and was correct. It felt then, and still feels, like a big step change on very hard problems. These are problems I would previously expect to take a month or longer to implement.

I have also noticed Fable can handle much more nuance when reasoning through writing and research, but that is harder to quantify.

What was the very complex feature?
Not OP, but to me it initially felt extremely proactive and energetic, just powering through roadblocks with ingenuity and enthusiasm. After it came back I was constantly getting refusals and downgrades for things Opus had been doing. I’ve written it off for my use cases and getting by just fine with Opus 4.8 and now 5.
Medium vs High? Why? From all the charts I've seen the performance jump is pretty large from med -> high (not as noticeable from high -> xhigh).
Medium or low supposedly prevents Opus 5 from overthinking:

https://xcancel.com/danshipper/status/2080700057892815114

https://cognition.com/frontiercode

Quality vs cost - medium is the sweet (perhaps better too!) spot.

That is just a single benchmark tho
my issue with frontier code is that it uses a model judge for quality whereas slop code bench forces a model to grapple with its own garbage code in order to receive a functionality reward
If I need something smarter I use Fable. Medium works well and is quick. Opus 5 medium feels much better to me than Opus 4.8 medium.
yeah someone will have to re-run this bench on various effort levels. unfortunately it is not cheap