Hacker News new | ask | show | jobs
by stavarotti 29 days ago
I'll be using it tonight but grudgingly so. Grudgingly because after July 7th, I'm not going to all of a sudden, start paying API prices (and maybe that's the problem) when I'm used to a subscription that gives me multiples in comparative value. Perhaps this is the fabled "token economics will come for everyone this year" that I've been reading about? In any case, I'll use the hell out of it to extract as much as I can, then back to the trusted partners Opus 4.6 and Sonnet 4.6 (for however long they remain available).
3 comments

Won't using it eat up the whole quota immediately forcing you to pay API prices anyway?
The token quota is completely unpredictable and changes month to month. Anthropic has a real penchant for riding the fine line of useful and dark patterns that make me want to write them off forever.
On the xhigh effort level (not ultracode!), and at the beginning of a new 5h session, I asked it to review a branch that has 300 lines changed/added, and went to grab a coffee. When I was back after a couple minutes, I saw that it decided to create a dynamic workflow with 60 something agents and hit session limits on my max plan.

When I'm subscribing to their plans, I have the expectation that I'll be able to get some reasonable amount of work done. These days, this expectation was already not being met for me with Opus, and Fable acting like this was the final nail in the coffin.

I cancelled my plan and I'm looking for alternatives.

For a period of time - then you go back to opus 4.8 or new sonnet 5.0 like some kind of AI pauper. Shine your shoes for some fable tokens g’vnor.
I am fully expecting the rollout of a Max 350 plan after July 7.
I locked my default model to opus 4.6 around the time of the nerfs. Such better results compared to 4.7+

That's enshittification for ya I guess

The claims of 4.6 or 4.7 being superior genuinely make me laugh. Adapt your workflow if needed and use the superior model instead of just kneejerk believing they actually enshittified a model with zero evidence except vibes on an undeterministic model output. Jesus.
4.6 was the last model that let you disable adaptive thinking and set max thinking token budget. I liked having that available, and still use it sometimes.
Your vibes are definitely better than his vibes.
What about all the benchmarks that show improvements in each generation?
Many of the improvements are the result of agentic loops and an emphasis on autonomy. Some of us don’t like that because the models go rogue and ignore design patterns, architecture, coding guidelines or other things that are important.

My friends and colleagues that like the agentic autonomy don’t care about the code, they feel like if it works it works and if an AI system is the only intelligence able to understand it that is ok.

I still want to be in the loop. They don’t.

the more agentic focused the better though?

sonnet 5 is very noticeably a much better model than any opus that ive touched

it actually does the things i want it to, and uses tools and triggers skills appropriately, vs trying to make stuff up

Agentic coding should absolutely care about all the things you listed.
Like another commenter said below, last Opus version to respect adaptive thinking and token budget flags was 4.6
It was quite clear 4.7 was a dumbed down high efficiency model they put out in a rush to handle the capacity issues they were having at the time. I've experienced myself substantial degradation on basic reasoning tasks, which were fixed in 4.8.
Bro, it's all vibes.

Models get dumber during the day and smarter during the night, I swear.

but I'm not willing to scientifically verify this, so I'm just going to go off of vibes- just like everyone seems to be doing with projects.

These vibes are pretty obvious even with casual use. Weekends are so much better.
In my case it that I'm tired and more likely to miss issues or mistakes. My idea of good enough is at a much lower level when it's 10pm and I'm about to knock off and go to bed in an hour.
Exactly.

The real step change I've seen lately is in the amount of complaining people are doing when their models aren't giving them what they want.

4.8 is much better than either of them as well.