Hacker News new | ask | show | jobs
by nullc 8 days ago
I think vibethinker is heavily overtrained on not attempting to solve open problems.

I had a fun time taking some open problems and disguising them algebraically so that vibethinker 3b would work on them. It managed to prove some interesting things that I didn't know and would be publishable, but for the fact that they already have been. :) (though hard to know if this was because it had been exposed to that knowledge even though it didn't reconize the hidden problem).

Under some maskings it would eventually figure out the problem was equivalent to an open problem then immediately shut down.

It also managed to make some false proofs for various things that duped some other more powerful models.

1 comments

IMO VibeThinker is the most interesting open model since the OG DeepSeek R1. The conventional wisdom has always been that specialization is not very helpful for LLMs, yet it outperforms models hundreds of times larger in its specialized area. It shows that there is a lot of fruit left to be picked, still out of reach but hanging low enough to be worth going back to the barn to fetch a ladder.

If I were a young Turk in this business, I'd drop everything else and figure out how VT3B is so ridiculously good at math.