Hacker News new | ask | show | jobs
by CamperBob2 9 days ago
Interestingly, even Qwen 3.6 27B was able to verify the solution, but I didn't get any glazing for discovering it. Instead, it thought that someone named Shestakov had already found a counterexample in 2004.

GLM 5.2 whiffed, it insisted the counterexample wasn't valid.

VibeThinker 3B also recognized that the counterexample was valid. But it kept trying to convince itself that it wasn't, over and over, since it's an "unsolved problem." Eventually it just answered "-2."

7 comments

I told Gemini Pro I woke up after having dreamt that polynomial and asked if it was related to the Jacobian conjecture. It spent some time thinking and referenced this tweet announcement, saying:

"If you truly dreamt about that specific polynomial, you might be mathematically clairvoyant."

In the rest of the answer, it maintained a cautious skepticism about my claim, saying:

"Here is exactly why the math world is currently scrambling to verify the polynomial you "dreamt" about."

I love how it put "dreamt" in quotes.

Must be that LLMs are just as suspicious of humans as some us are of them.
> Matches! This is bizarre. A Jacobian counterexample has been sitting here in a prompt? Wait... is this map a known "fake" counterexample from the literature? Many mathematicians have tried and failed. This specific map might come from a paper or a forum where it was proposed and then debunked. Or... is it actually correct?

Gemma's having trouble accepting it too. A solution?! At this time of year? At this time of day? In this part of the country? Localized entirely within my own prompt?

Inconceivable!
Yes!

Can I see it?

No.

Current LLMs behave very counterproductively around unsolved problems, especially if they learned that humans consider them difficult. This has many straight up preventing themselves from attempting anything...
Haha, I’ve noticed this as well. It’s like they psyche themselves out about how famous the problem is the same way humans do. I gave one Collatz in disguise, and it was finding all sorts of interesting things (but nothing worth a paper) until it realized the problem was Collatz, at which point it just proceeded to find a bunch of reasons why nothing would work from that point on.
I think vibethinker is heavily overtrained on not attempting to solve open problems.

I had a fun time taking some open problems and disguising them algebraically so that vibethinker 3b would work on them. It managed to prove some interesting things that I didn't know and would be publishable, but for the fact that they already have been. :) (though hard to know if this was because it had been exposed to that knowledge even though it didn't reconize the hidden problem).

Under some maskings it would eventually figure out the problem was equivalent to an open problem then immediately shut down.

It also managed to make some false proofs for various things that duped some other more powerful models.

IMO VibeThinker is the most interesting open model since the OG DeepSeek R1. The conventional wisdom has always been that specialization is not very helpful for LLMs, yet it outperforms models hundreds of times larger in its specialized area. It shows that there is a lot of fruit left to be picked, still out of reach but hanging low enough to be worth going back to the barn to fetch a ladder.

If I were a young Turk in this business, I'd drop everything else and figure out how VT3B is so ridiculously good at math.

More anecdata: When I just tried Qwen 3.6 27B (Q6_K_XL) it (ultimately, after a lot of going back and forth) claimed it was not a counterexample and claimed the Jacobian wasn't constant (which I'm guessing is incorrect). It also mentioned a whole bunch of names it attributed the example to, in its thinking trace.
I only asked once, so it might well be inconsistent. I did try asking GLM 5.2 NVFP4 several times, and it returned consistent wrong answers at both thinking and max-thinking levels.

For Qwen 27B, I have better luck with a Heretic-derived 8-bit quant than I did when I was trying to run the various smaller GGUFs.

Connecting it to a Coding Agent seems much better.

I connected DeepSeek in OpenCode and told it that I dreamed of this counterexample. It called SymPy tools to verify it, said my dream was "surprisingly accurate", and suggested consulting an expert in algebraic sets for independent verification.

Also interestingly, I told it: "DeepSeek told me the Jacobian conjecture is false, and gave me a counterexample. How should I treat this?"

He immediately told me that this DeepSeek was talking nonsense. Someone who can give a real counterexample "would not be a bot from an AI company, but a Fields Medal winner."

> that someone named Shestakov had already found a counterexample in 2004.

Qwen has the sprit of a grad student

It must mean Chestikoff. He dedicated his life to intellectual pursuits due to chronic ill health.
There's a certain (unfair) stereotype of Russian mathematicians that this fits.