If you say "find some unsolved graph theory problem and counterexample for it" and LLM actually does it, is it really you that solved the problem? That's the difference vs other tools.
It's not, it's basically something that happened. One unsolved problem was solved with the only human input being telling LLM to work harder and to actually solve it after few failed attempts.
Or are you saying that’s what happened in this case. Because that’s not the way I understand it.