|
|
|
|
|
by Legend2440
11 days ago
|
|
I don't agree with your definition. The point of reasoning models is that some tasks fundamentally require a certain number of serial steps. Base models are limited to learning parallel algorithms because of their parallel training, and so struggle on inherently-serial tasks like solving logic puzzles. The advancement from reasoning is that it allows LLMs to learn a broader class of algorithms. |
|
The problem with "reasoning" is that for the most part it happens outside of training e.g. as CoT. The base model knows what it knows, it knows what it's trained on, and it can't really go much farther than that.
Now, starting with o3, "reasoning" models are probably (who knows exactly) trained on traces of reasoning, either from automated systems or from human experts, but that doesn't mean they learn any kind of algorithm, just that they learn to reproduce the behaviour of "algorithms" (or of reasoning humans). That can improve performance on certain kinds of task (the ones in the training set) up to a point, but you're not going to make a tiny model perform like one ten times larger just by that.
On the other hand, I do think that LLMs can be made smaller without losing a commensurable amount of performance, like e.g. the Tiny Recursion Model (TRM) which did OK at ARC 1 (45% I think). But then you lose a lot of functionality also. Basically the larger models probably have more parameters than they really need and that's something the industry seems to have realised, but you still need huge parameter counts to reach top performance anyway.
And then there's the training tokens, which aren't getting any fewer.