Hacker News new | ask | show | jobs
by amluto 18 days ago
No, the comment is right. The prompt had GPT-5.6 reviewing the proof, and the result, unsurprisingly, survives review by GPT-5.6.
1 comments

Given a new context, why couldn't the same model have a decent shot at reviewing some results? It's not like they identify whether this output is from them and then go "yeah correct", that's not how they work.
It’s the other way around. The prompt instructed a GPT-5.6 agent to try to make a proof that would survive review by a GPT-5.6 subagent. If there were some defect that would cause the reviewer subagent to accept an incorrect proof, then one might imagine that someone else asking the same model to review the same proof would give the same result. And the proof generation process might even be biased to find such an incorrect proof.