Hopefully it stands for AC Generation Improvements. If it prioritizes income it will bleed the planet dry. It needs to solve how expensive our cost is on the planet first or its entire existence was a mistake.
The definition has already been stretched to not fit the previous models. There is no meaningful, static definition that significantly predates current capabilities.
There's a reason why ai xrisk doomers had to come up with the term ASI.
I would seriously suggest that everyone take a look at the wikipedia page for AGI from the month before ChatGPT was released, compare it to the current version, and not come to that conclusion.
The first sentence is “understand or learn any intellectual task that a human can.” Whatever you think of the benefits of LLMs, they don’t understand and they can only learn during the training period and with very minor adjustments in post training. So, no I don’t think any of these models are generally intelligent.
I have not seen any instance of this frequently-made assertion which is at all justified. It seems to rely on a definition of "understand" which is more about spirituality than actual observable evidence (they clearly can comprehend even complex tasks well enough to execute on them, and if you won't call that "understanding", you're playing word games rather than stating an objective fact).
Likewise, agents can literally come to a greater understanding of a problem through trial and error, and there are plenty of mechanisms to retain that knowledge. If you don't want to call that "learning", you're just making a choice to define it in a way more restrictive than how we use it for humans, and intentionally making communication more difficult.
Consider the number of humans I've seen make statements that fit that description (about AI, no less!), I don't think that's a strong argument against it.
Agents are always combining the same underlying weights to their inputs, relying on the same maps of semi-semantic space and the relationships between those that it was leaning towards at training time. The fact that it’s successful in making lots of people have an Eliza effect doesn’t make it understand something. It’s simulating understanding based on an enormous corpus of text, much of which is people working through things or sharing an understanding of something. Unless you believe that all intellectual activity is about finding the space between words you shouldn’t believe LLMs have any chance at understanding anything.
It seems to rely on a definition of "understand" which is more about spirituality than actual observable evidence
"Understanding" has enough philosophical leeway in its use to allow at least the possibility of sentience as a prerequisite.
This is where the discussion about LLM capabilities becomes genuinely difficult, and dismissing that difficulty as "word games" or "spirituality vs evidence" is not helpful.
Considering that "sentience" has enough "philosophical leeway" that it's just as reasonable to assert that LLMs are sentient (and at extremes, that they have been sentient for years) -- especially if we are, as you suggest, supposed to include any philosophically possible definition -- I don't think that's a meaningful rebuttal. If no one can agree on whether it's sentient, it's bad faith to choose a fringe definition that hands off its definition to such a nebulous term.
In fact, I'd argue that statements about what "is" and "is not" sentient relies on even more spirituality and word games for anything that isn't a terran tetrapod.
For a meaningful -- "helpful" -- discussion on such things, one has to assume that everyone is choosing a definition which is closer to the median usage and relies on not being totally subjective. Furthermore, given the breadth of options, it should be assumed to be a definition which allows which permits the form of the question to be meaningful, rather than begging the question -- if your definition is tautological enough that non-biological entities can't have understanding, you're just expressing dogma rather than having a discussion.
Anything else is bad faith, or assuming bad faith on the part of the participants.
AA-Briefcase is a new benchmark for testing models on realistic knowledge work tasks in complex projects built by industry experts. Models are evaluated on multi-week knowledge work projects, each with many linked tasks and thousands of input source files. AA-Briefcase combines rubric and pairwise grading to evaluate verifiable task success, analytical quality, and presentation quality, giving a holistic view of overall agentic capability in knowledge work.
Tasks with many messy input files, conflicting information, and complex deliverables remain difficult for all models. Under a strict all-or-nothing grading scheme per task, Claude Fable 5 leads overall, but achieves a perfect task score on only 3% of tasks. On 31 of 91 tasks, no model scores above 50%.
AA-Briefcase is a new benchmark for testing models on realistic knowledge work tasks in complex projects built by industry experts. Models are evaluated on multi-week knowledge work projects, each with many linked tasks and thousands of input source files. AA-Briefcase combines rubric and pairwise grading to evaluate verifiable task success, analytical quality, and presentation quality, giving a holistic view of overall agentic capability in knowledge work.
Tasks with many messy input files, conflicting information, and complex deliverables remain difficult for all models. Under a strict all-or-nothing grading scheme per task, Claude Fable 5 leads overall, but achieves a perfect task score on only 3% of tasks. On 31 of 91 tasks, no model scores above 50%.
I think you're speeding past the word "average" in the sentence. I'd argue that current frontier models already exceed the abilities of average humans across the majority of tasks you can do on a computer, although you might be able to argue that they tend to be a bit slower?
That latter part is debatable though - have you seen a non-technical person try to figure out something new on a computer?
AA-Briefcase is a new benchmark for testing models on realistic knowledge work tasks in complex projects built by industry experts. Models are evaluated on multi-week knowledge work projects, each with many linked tasks and thousands of input source files. AA-Briefcase combines rubric and pairwise grading to evaluate verifiable task success, analytical quality, and presentation quality, giving a holistic view of overall agentic capability in knowledge work.
Tasks with many messy input files, conflicting information, and complex deliverables remain difficult for all models. Under a strict all-or-nothing grading scheme per task, Claude Fable 5 leads overall, but achieves a perfect task score on only 3% of tasks. On 31 of 91 tasks, no model scores above 50%.
" I'd argue that current frontier models already exceed the abilities of average humans " for things that fit in their context window sure but LLMs can't learn over time the way humans can. One example is LLMs are very good at writing a few thousands line of code but they absolutely cannot write coherent million line codebases. By average human I meant the average skill level for the job. AGI would need to be able to pass a interview and get hired and the perform well enough to not get fired.
Yeah it's not true that for every job, it is better than median worker of that job. But it is conceivable that for almost all jobs it is already better than the median human (not just workers of that job).
Continual Learning? Why is this even a question? Isn’t it a well-known glaring issue with the current models? They cannot learn/adapt to new skills (in any permanent sense) once they are deployed.