Hacker News new | ask | show | jobs
by davidpapermill 12 days ago
> It would've been nice to reserve "AI" for superior human-like intelligence capable of genuine common sense and reasoning.

What would a frontier API have to be able to do to satisfy you?

5 comments

Maybe just a very rigorous version of the Turing test? Modern LLMs can superficially simulate conversation but it's trivial to force them into revealing their non-human like intelligence.

They've been "patched" since but all models fail basic tests like "Should I walk or drive to the car wash which is 100 feet away" by recommending you walk.

So you'd just ask questions that require theory of mind, abstract and common sense reasoning, causal inference, learning novel rules, transferring knowledge novel situations, recognizing ambiguity, etc.

If an alien lands on Earth and learns English, would you deem it non-intelligent if you can tell it apart from a human in conversation?

I think we should consider slime mold intelligent, and realise that it's a spectrum. Path finding is AI. There are probably forms of intelligence we have yet to discover.

If an alien landed we could decide whether it seems to have a human-like intelligence or not. It could be incredibly intelligent but very non-human-like.
Can you give me one example that works on Claude right now?

I'm never sure whether this indicates "no reasoning present" or you've just hit an odd behaviour in the AI such that its reasoning fails. For example, you present a problem in a way that's dissimilar to the way problems are presented in its training set. That doesn't mean it's not reasoning, just it can only reason correctly in some circumstances.

The models are continually patched with training and post-training. All you have to do is find an area they haven't patched yet, and they'll be just as stupid. I run into deep technical examples every day where they fail in the most basic ways no human ever would.

I'm pretty sure most people building these models would admit they don't operate as human-like intelligences? It's baffling that anyone thinks they are.

Yes I agree, they’re an alien kind of intelligence.

But that doesn’t mean they don’t reason.

I get what you're saying but this is kind of a semantic game.

These LLM models/agents absolutely do not reason in the sense that humans do, so you're quietly redefining the word.

You can say of course decide to call them an "alien kind of intelligence" that "reasons" but you could just as reasonably say that calculators are an "alien" kind of intelligence that "reasons" about math differently than us.

And what stops what AIs do from being "reasoning"? What's the elusive magic fairy dust of reasoning that humans put into their napkin notes, but AIs neglect to put into their chain of thought scratchpads?

Do you have a RealReasoningBenchmark, perhaps, that can reliably tell apart that fake mass produced token-flavored AI reasoning from the real, organic, 100% natural human reasoning?

If you had access to a bunch identical copies of me that couldn't communicate with each other, you'd be able to find many questions I would give stupid answers to. I suspect I'd come out of it looking worse than an LLM.
How old does a child need to be before you think they have "human like intelligence"?
I would walk
Me too. At least it doesn't say I need to wash my car.
To answer for OP:

We are now calling text and image generators "intelligent" in the same way a spell checker is intelligent.

Whatever it's become, "AI" research started as a way to study digital neurology, or how to digitize a mind, not just how to generate data.

The Turing Test should have had a caveat, it needs to fool a, "non-stupid" person, and we still have not gotten even close to passing that version.

What exactly would a 'non-stupid' person do to catch the latest models on a Turing Test? Aside from being aware of AI 'tells' like em-dashes.
If this was true. I would repalce myself with ai that pretends to be me on slack.

my coworkers would know almost immediately if i did that.

> my coworkers would know almost immediately if i did that.

The same would happen if you were replaced by any random human.

But immitating others is about the only thing genAI does. Sometimes "others" is a 'programmer', sometimes "others" is an 'artist', but regardless, it still does it poorly.
that would be a silly test then
It sounds kind of like you made up a silly test then.
Not OP, but I'd settle for something that actually learns, instead of being a static pile of linear algebra. Pretending it learns because you change the input (context) doesn't count.
Can it produce a chart topping album if its given all the tools and the prompt "produce chart topping album" .

you might say almost no humans can do tht either but some human can but no ai can.

strawberry