Hacker News new | ask | show | jobs
by canada_dry 20 days ago
The reason I'm getting LLM burnout is from dealing with the obvious neutering and opaque downgrading of all the top models.

Prior to the last 12mos AI companies were hell bent on squeezing out the best results from mediocre models.

But... now that the top models have progressed, those same AI companies have switched their efforts into reducing the computation (cost of a producing a result) as much as possible without being too obvious.

What was an exponential slope in the quality of results over the last 36 months has now nearly flat lined.

Addendum: IMHO results have 'flat lined' not because the models aren't much more capable than a year ago, but because conserving the enormous processing cost (of an over subscribed user base) supersedes the goal of following the user's explicit instructions (e.g. especially if that means more processing cost) to generate the best results.

4 comments

I feel the same way about consumer AI tools now. Gemini and ChatGPT have been abysmal lately. They can no longer be relied on to do multi-turn searching and thinking.

Before, they could stay in thinking mode for more than 7 minutes. For example, "find a source for this claim" would search, analyze, and self-adjust the query. Nowadays, even if I push for it, I cannot make these tools work for more than 30 seconds before they give generic answers, even in "Pro" mode.

At one point we considered adding artifical delay to responses because irrational users dont trust something that finishes fast, even if its the same quality.

How empirical are your comparisons of new and old outputs?

Counterpoint: Reels apparently are still addictive
You sure about that? Maybe it is reality hitting expectations after the initial “holy shit” wears off
Smart people have been falling into this trap as long as LLMs have hit production. Supposedly early internal versions of GPT-4 had "sparks of AGI" but the public version was "dumbed down for safety"

https://www.youtube.com/watch?v=qbIk7-JPB2c

I'd bet more on this personally.
This seems hilariously, extremely revisionist.

Hell, the Opus 4.5 moment was only last November, and that was when agentic coding and most coding CLI tools became truly first class options. That's a wild paradigm shift. Hell, GPT-5 wasn't even out (that's August of last year). Most people were using 4o. Their current offerings are wildly better for coding than 4o was.

I generally don't agree with the original commenter here. I think many of the complaints about model regressions are the result of increased usage and increased scrutiny revealing gaps there were there all the time. I've been more critical than most of the output quality since my initial "wow" moment was pretty early - GPT 3.5 API - and the results then were extremely obviously not production ready. But, keeping that level of scrutiny through my usage, I haven't seen the falloff that people who don't look at the output every time claim to see.

But that's also let me use "agent" stuff longer, I guess? The better you were at knowing what you wanted and how to ask for it, the less of an inflection point that you got from Opus 4.5 or GPT 5.

Some of the highest-time-saved-for-max-ROI agentic problems I've solved to date were in September and October of last year with Claude or Cursor.

Cannot relate, my expectations might just gone up, because when I compare what I was producing with agents a year ago vs now, it's night and day.