Hacker News new | ask | show | jobs
by YeGoblynQueenne 21 days ago
>> What’s left?

For example, there's all the problems that the same off-the-shelf model hasn't solved despite OpenAI running it for many hours on them. Don't forget you're only seeing the results of successful runs.

We can estimate that those unsolved problems must number in the dozens, or even hundreds, given the amount of time that passed since the last announcement of a solution to an interesting problem by an OpenAI model: i.e. the unit distance problem which was announced solved in 20 May this year. That's a couple of months, yes? We can be fairly certain that OpenAI have been trying to solve other problems all this time, first because they are hell bent on demonstrating that their models can do maths and second because we just got another result, but it took that long. They were obviously not twiddling their thumbs all this time.

So if OpenAI are running their model on a single proble for eight hours at a time (according to the prompt they released) they could be easily have run a few hundred instances of their model on the same number of open problems 156 times for each instance (53 days since 20 May, with a model running in three eight-hour sessions per 24 hour day). I mean the only restriction is the cost they're willing to pay for the inference.

So yeah, there's a lot left to do still, don't worry.

1 comments

The 2 most notable/interesting solutions have come from Open AI directly, but most of the 'LLM solves open problem' category didn't and has come from 3rd parties doing their own thing with publicly available models. I don't see why one would assume they're running models on hundreds of problems. Most likely they have a few problems they especially care about that they run on.
What, Erdős problems? It's hard to see how anyone except mathematicians, and then again only a few communities of mathematicians, especially care about those.

Remember back in the day when Deep Mind made AlphaGo? That made huge waves for two reasons: one, it was very well understood by AI researchers that beating expert humans at Go was very hard; and, two, that Go is a game of great cultural significance to literally billions of people outside of academia, albeit mainly in SE Asia.

Now, Erdős? I'm a computer scientist and I had to look up the planar unit distance problem when I heard about it because I had no idea what that was. Because it never comes up in the literature I read. I'm not saying it's not interesting, but it does seem a bit... random? That they started with Erdős problems? I'd have gone for a Millennium Prize problem, first. P vs NP, Riemann, Navier Stokes, those are heavy-weight results that would establish AI as the de facto approach to mathematics for the foreseeable future. Even I would find it hard to raise an objection (imagine that). Erdős can take a number, compared to all that.

I'm not saying they're choosing problems at random, mind. But it does seem like we're only seeing the tip of the iceberg, with respect to what the AI companies are doing internally. That shouldn't be a surprise. That's exactly how research works in general, both in academia and in industry. There is a clear survivorship bias and we only ever get to see the positive results, never the negative ones.

>> The 2 most notable/interesting solutions have come from Open AI directly, but most of the 'LLM solves open problem' category didn't and has come from 3rd parties doing their own thing with publicly available models. I don't see why one would assume they're running models on hundreds of problems.

Actually, that's a good point but it's in support of my contention. If there are random people in the community running LLMs on their own, favourite, maths problems, we should be seeing many more of those solved and much more often, provided LLMs were really as good at maths as OpenAI et al want them to be. There must be literally thousands of mathematicians trying to use LLMs to solve this or that problem that is famous in their community. Where are all those spectacular results?

Or, to abuse Fermi's question, where is everybody?

>What, Erdős problems? It's hard to see how anyone except mathematicians, and then again only a few communities of mathematicians, especially care about those.

Erdős problems vary enormously in difficulty and significance. The fact that a problem is obscure to non-mathematicians does not make its solution unimportant. There are also many major open problems in computer science that most laypeople have never heard of.

>That they started with Erdős problems? I'd have gone for a Millennium Prize problem, first. P vs NP, Riemann, Navier Stokes, those are heavy-weight results that would establish AI as the de facto approach to mathematics for the foreseeable future.

Solving the biggest, most famous open problems would "establish AI as the de facto approach to mathematics for the foreseeable future"? It would do a lot more than that.

>That's exactly how research works in general, both in academia and in industry.

Then what exactly is the objection? Research normally produces many failures and incremental results before a major success. Do you apply this survivorship-bias criticism to every published mathematical result, or only when a machine contributed to it?

>Actually, that's a good point but it's in support of my contention.

If they're running as many problems as constantly as you imagine then they did not miss all that. So either they're just not sharing it us which contends with your "desperate to demonstrate mathematical competence" or they have their eyes on a more curated set.

>Or, to abuse Fermi's question, where is everybody?

If we had as many verified alien encounters as LLM contributions to open problems, nobody would be invoking the Fermi Paradox. Where is everyone? Right here.

>> Then what exactly is the objection? Research normally produces many failures and incremental results before a major success. Do you apply this survivorship-bias criticism to every published mathematical result, or only when a machine contributed to it?

Yes I do. Not mathematical results specifically but generally research results. I might even have articulated that criticism on HN. I don't know if I could search for it easily though.

>> If they're running as many problems as constantly as you imagine then they did not miss all that. So either they're just not sharing it us which contends with your "desperate to demonstrate mathematical competence" or they have their eyes on a more curated set.

The people desperate to demonstrate mathematical competence are the AI companies. The people discussed in the part of my comment you quote are "random people" by which I meant the "3d parties" in your original comment.

>Yes I do. Not mathematical results specifically but generally research results. I might even have articulated that criticism on HN. I don't know if I could search for it easily though.

Okay Fair, but then this is a general research issue and not really a Open AI issue.

>The people desperate to demonstrate mathematical competence are the AI companies. The people discussed in the part of my comment you quote are "random people" by which I meant the "3d parties" in your original comment.

I'm not sure you got the point i was making. The point there was that if Open AI were running as many problems as frequently as you imagine they are then those 3rd party results should have been achieved by them, and if they're really so desperate to tell us how good the model is for math then why didn't they tell us ? Why have they only announced 2 results when they could have announced near a dozen by now ? Sure these 2 are in a class of their own, but some of the others are genuinely impressive in their own right and would certainly help that narrative you're talking about.

Either they're just not telling us and aren't as desperate as you imagine, or they're simply only interested/running models in a relatively few set of problems.

Or they're not solving as many problems as they'd like, despite trying.