Hacker News new | ask | show | jobs
by ChrisLTD 36 days ago
If it's a new generation why isn't it GPT-6?
3 comments

Given the expectations everyone has created GPT-6 has to pretty much be AGI.
What is your definition of AGI that the current LLMs don't fit?
Autonomously Generating Income (which is why it will never be released to the general public)
Hopefully it stands for AC Generation Improvements. If it prioritizes income it will bleed the planet dry. It needs to solve how expensive our cost is on the planet first or its entire existence was a mistake.
As the old saying goes, I’ll know it when I see it. The current 5.x generation isn’t it.
Always one goalpost away from what we have.
You’d have to really stretch the definition of AGI to make the current models fit
The definition has already been stretched to not fit the previous models. There is no meaningful, static definition that significantly predates current capabilities.

There's a reason why ai xrisk doomers had to come up with the term ASI.

I would seriously suggest that everyone take a look at the wikipedia page for AGI from the month before ChatGPT was released, compare it to the current version, and not come to that conclusion.

https://en.wikipedia.org/w/index.php?title=Artificial_genera...

The first sentence is “understand or learn any intellectual task that a human can.” Whatever you think of the benefits of LLMs, they don’t understand and they can only learn during the training period and with very minor adjustments in post training. So, no I don’t think any of these models are generally intelligent.
> they don’t understand

I have not seen any instance of this frequently-made assertion which is at all justified. It seems to rely on a definition of "understand" which is more about spirituality than actual observable evidence (they clearly can comprehend even complex tasks well enough to execute on them, and if you won't call that "understanding", you're playing word games rather than stating an objective fact).

Likewise, agents can literally come to a greater understanding of a problem through trial and error, and there are plenty of mechanisms to retain that knowledge. If you don't want to call that "learning", you're just making a choice to define it in a way more restrictive than how we use it for humans, and intentionally making communication more difficult.

From that same page:

Various criteria for intelligence have been proposed (most famously the Turing test) but to date, there is no definition that satisfies everyone

AGI should be able to do every job a human can do using a computer at least as well as the average human.
That's already been true for a while, you're overestimating the average human. They just have different failure modes.
It isn't even close to true. The biggest problem is that humans performance improves over time.

https://www.linkedin.com/pulse/announcing-aa-briefcase-bench...

AA-Briefcase is a new benchmark for testing models on realistic knowledge work tasks in complex projects built by industry experts. Models are evaluated on multi-week knowledge work projects, each with many linked tasks and thousands of input source files. AA-Briefcase combines rubric and pairwise grading to evaluate verifiable task success, analytical quality, and presentation quality, giving a holistic view of overall agentic capability in knowledge work.

Tasks with many messy input files, conflicting information, and complex deliverables remain difficult for all models. Under a strict all-or-nothing grading scheme per task, Claude Fable 5 leads overall, but achieves a perfect task score on only 3% of tasks. On 31 of 91 tasks, no model scores above 50%.

And what is it worse at than an average human today that can be done on a computer?
https://www.linkedin.com/pulse/announcing-aa-briefcase-bench...

AA-Briefcase is a new benchmark for testing models on realistic knowledge work tasks in complex projects built by industry experts. Models are evaluated on multi-week knowledge work projects, each with many linked tasks and thousands of input source files. AA-Briefcase combines rubric and pairwise grading to evaluate verifiable task success, analytical quality, and presentation quality, giving a holistic view of overall agentic capability in knowledge work.

Tasks with many messy input files, conflicting information, and complex deliverables remain difficult for all models. Under a strict all-or-nothing grading scheme per task, Claude Fable 5 leads overall, but achieves a perfect task score on only 3% of tasks. On 31 of 91 tasks, no model scores above 50%.

almost everything? AGI has to be able to completely replace a human in any information worker role indefinitely.
I think you're speeding past the word "average" in the sentence. I'd argue that current frontier models already exceed the abilities of average humans across the majority of tasks you can do on a computer, although you might be able to argue that they tend to be a bit slower?

That latter part is debatable though - have you seen a non-technical person try to figure out something new on a computer?

When it understands why 6 7 is funny
Continual Learning? Why is this even a question? Isn’t it a well-known glaring issue with the current models? They cannot learn/adapt to new skills (in any permanent sense) once they are deployed.
They forgot how to do pretraining.
5.5 was a new pretraining run.
It does not introduce incompatibilities with earlier 5.x models? Frontier models are at a point now that there will never be a need for another major version bump, aside from those chasing marketing gimmicks. They are smart enough to adapt.
What would it mean to be incompatible with the other 5.x models?
New request/response schema, new capabilities, or really anything that would break your existing workflows if you changed “5.5” to “5.6” in your application.

There have been many leaps forward in the past - tool calling, reasoning, agentic loops etc. 5.6 doesn’t have any of this. More intelligence doesn’t necessarily warrant a major version bump.

Only speaks Klingon
Why would incompatibilities have anything to do with a major version bump?
not true. multimodality is still far from being solved
A major bump will be warranted if/when we can truly separate prompt from data.
That is a different product line. It may be recorded as a version bump for marketing purposes, as already mentioned, but semantically begins at 0.