Hacker News new | ask | show | jobs
by YeGoblynQueenne 14 days ago
Without human verification, an LLM can generate correct or incorrect proofs but it can't tell the difference. A human is necessary to be able to tell one from the other.

Saying that's a solution "done autonomously by an automated AI pipeline" is like saying that a self driving car that can only take you to the nearest train station after which you have to ride the rain to where you're going is "autonomously" driving you to your destination. Which is exaggerating the autonomy of the system, rather.

1 comments

The automated AI pipeline also had an automatic grading model to try to reduce false positives.

But anyway, my point was that in that case the prompt involved was indeed pretty much "hey, ChatGPT, solve an unsolved problem, thanks."

Yes, but there's no way to tell how many times the prompt was tried before a proof was found; and that only stopped when a human said it could stop.