Hacker News new | ask | show | jobs
by IanCal 4 days ago
> I've worked with these systems for four years now and they have not meaningfully improved in that time frame.

Not meaningfully improved?! Four years ago was gpt *3.5*! ChatGPT hadn’t been released!

1 comments

Yes! Impressive, isn't it? I see how it has improved for some minor points, that the big models can cover more finetuning ground, but my big gripes are still the same - you could do the same back then with multiple models and more targeted finetuning.
> you could do the same back then with multiple models and more targeted finetuning

Definitely not, lol.

I have no idea how you could try and hold this argument without being facetious.
Am I missing something that my original points no longer hold for their products? Did it get meaningfully solved? Are your experiences flawless on that front?
What do you hope to achieve here? Posting absurd and ridiculous things and then continuing like if someone would take you seriously after that. Keep going I guess.
You're moving the goalposts. Initially it was "no meaningful improvement" and now suddenly it has morphed into "they're not flawless".

I'm pretty sure you're just baiting for engagement though so well done, ya got me.

This is obviously not true to all of us here…
> you could do the same back then with multiple models and more targeted finetuning

I mean, come on, this is just not true. You could not achieve anything like what you can with modern agentic coding with Fable / 5.6 Sol from any combination or configuration of GPT 3.5 era models.

It's like saying that a teenager isn't an intellectually meaningful improvement over a toddler.

Sure they're both still fundamentally flawed humans prone to cognitive error, but one is clearly more likely to hit the mark than the other when assigned a task.

The only thing impressive is how wrong you are. LLMs have improved by an absolutely incredible amount in the last 4 years.
Utter nonsense.

There’s no way you could get models as smart by fine tuning. I couldn’t throw a problem like “build a pokemon database with UI to teach my son sql” and get a working system, nice ui, tests (which it iterated on) examples and explanations in one shot.

There weren’t thinking tokens. Maths is now dramatically better, making actual contributions when before they were mostly mocked for making extremely basic errors. Smearing is also something say is very rare in frontier models.

If you think they have barely changed you’ve either forgotten what they were like or not used them more recently, or you’re just being obtuse.

My company had such a system four years ago, for internal work, somewhat more limited in scope (one language). What you are seeing as the frontier is not necessarily the best you can have - just because people don't try to push it to market as a product doesn't mean it's not there.

Edit: we do have a system that uses LLM and fixes the above issues largely (tracking of state, calculations and objects, still flawed in finer details). No, we don't sell, it's experimental fun and not really ready in terms of setup/ux/etc.

It codes really well for our case though.

> you could do the same back then with multiple models and more targeted finetuning

Are you one of those anonymous billionaires as if you did this a few years ago, you would've been famous and rich.

OP is delusional or deliberately optuse. I work in the space and stare down these systems 12h/day, and saying the systems haven't meaningfully improved is ludicrous.
OP is largely pissed with what OAI/Antrophic are trying to sell as meaningful improvements and the market-bending money they ask for it. I work in the space and we trained LLM models on conceptual tokens, not language tokens, for example. See Symbolic AI and all the attempts at hybrid models.

Also, uh, fame and riches are not really my thing. Middle income is fine. My mistake was speaking up here because I got carelessly annoyed because I have skin in the game, research-wise. I'm sorry for that.

> are trying to sell as meaningful improvements

There may not be "core" improvements (structural reliability) but there are "emergent" improvements (apparent intelligence). Already the IQ tests from Maxim Lott ( trackingai.org ) show a progressive sliding towards the right side of the curve - which btw translates to a very much non-secondary decline in the user's frustration (and progresses with an increase of usability).

The more they work on it, the more probable the jump becomes - e.g. to achieve the Large Conceptual Models you say you worked on.

I still enjoy the symbolic ai space. Any examples of interesting progress there?
You're welcome to speak up, but saying models haven't meaningfully improved in three years, when the most recent models are solving top-level frontier maths problems, is indeed going to strike most people as weird. I still have no idea what you mean.
I think he means that, while results are there, they are mysteriously emergent, since an analysis of the process reveals it can be faulty.