|
|
|
|
|
by root_axis
23 days ago
|
|
The difference is a lot more than just throwing scale at it, pretty much everything useful comes from an evolving landscape of post-training techniques. Of course, param count and context length are also important because they increase the model's overall fidelity, but a base model without SFT, RHLF etc is effectively useless. |
|