|
|
|
|
|
by BoorishBears
317 days ago
|
|
No, I've deployed a lot of open weight models and the gap between closed source is there even at larger sizes. I'm running a 400B parameter model at FP8 and it still took a lot of post-training to get an even somewhat comparable performance - I think a lot of people implicitly bake in some grace because the models are open weights, and that's not unreasonable because of the flexibility... but in terms of raw performance it's not even close. GPT-3.5 has better world knowledge than some 70B models, and a few even larger. |
|
Without constantly refreshing the underlying LLM and the expert system layer, these models would be outdated in months. Language and underlying reality would shift from under their representations and they would rot quick.
That's my reasoning for considering this a bubble. There has been zero indication that the R&D can be frozen. They are stuck burning increasing amouts of cash for as long as they want these models to be relevant and useful.