Hacker News new | ask | show | jobs
by suprjami 9 days ago
You had me up until you equated LLM usage to unreliable human delegation.

No human being is going to genuinely suggest to glue cheese onto a pizza to stop it sliding off.

These things aren't people. Stop anthropomorphizing a computer program. They'll make all the mistakes humans can make, as well as a new class of mistakes because they're just fancy next word prediction with no lived experience.

3 comments

Indeed.

Just to add - I’m working on something very novel and LLM’s are absolutely useless at them. It’s actually comical to see the outputs - when you work on novel stuff you quickly see what these models actually are.

Apparently if you show one the recently found counterexample to the Jacobian conjecture, it spins in circles double-checking it because it knows the conjecture is true.
And finally gives up admitting it is a counterexample. Even saying "holy shit". That's what any human expert would do: presented with an extremely improbable fact, it would verify it, then in disbelief verify it again, then change method, then try yet another one, checking its understanding of the problem itself... and finally concede. Because this is what intelligence looks like- understanding the problem, its relevance, its context, self-doubting if the solution looks too simple to be true, double-checking all calculations and trying different approaches, etc.
I think Haiku still does this with the seahorse emoji. Doesn't feel like it's finding contradictory evidence, though, just feels like it's trying both options in case one works.
Precisely - except you can get it to do the same thing about the flat earth or faked moon landings or the unity of all being as will or just about anything else.
> No human being is going to genuinely suggest to glue cheese onto a pizza to stop it sliding off.

That mistake was made more than two years ago by whatever experimental version of Google's AI overview (which has strict performance requirement- i.e. needs to answer within a split-second) was up at the time. In other words, it was a very small and primitive system specialised in spitting answers as quickly as possible without a second thought. I hope you realise that basing your assessment of what LLMs can or can't do on that example is not much better than suggesting to put glue on pizza. A mistake that, if we were to adopt your reasoning, would in turn set a hard limit to the analytic skills of all humanity.

It's a clear illustration that these systems are not oracles. No matter how much they've improved since then, we should assume they are still capable of making hilariously big mistakes. Only now the obvious mistakes are fixed and the big mistakes will be a lot more subtle and a lot harder to catch.
> It's a clear illustration that these systems are not oracles.

Of course they are not oracles. Oracles belong to magic, these things are real.

I really can't understand how people seem to expect "intelligence" to be deterministic and infallible and uniform across domains when all examples we have of it (that is, us) are anything but.

> I really can't understand how people seem to expect "intelligence" to be deterministic and infallible

We don't really expect this of intelligence

We do sort of expect this of machines

Most people in society use machines and computers expecting very deterministic behavior, within some constraints. A saw cuts wood as long as your blade isn't dull. A computer shows your emails as long as the power and Internet is not out.

Comparably Unconstrained computer AI systems are still extremely new and there is no social norms or expectations around it yet

> I really can't understand how people seem to expect "intelligence" to be deterministic and infallible and uniform across domains when all examples we have of it (that is, us) are anything but.

I’m sure there are people who think putting glue on a pizza is acceptable but we’re not hiring them to bake pizzas for us.

When I search and get an AI summary it’s either correct or useless. I don’t care how fallible humans are, I’m not asking them to summarize my Google results.

So you never go to a doctor because (being human and therefore fallible) they might give you wrong information?
Nice try with the gaslighting. Google's AI "summaries" get things wrong all the time even now.
Hey, original author here.

> But fundamentally any metaphor will be inadequate, LLMs are too weird. It doesn't make sense to compare a transformer-based deep neural network that autoregressively generates language tokens based on their embedding in a high-dimensional semantic space to really anything that's ever existed.

You might not have read the full post ha