Hacker News new | ask | show | jobs
by Yopolo 10 hours ago
Im pretty sure they know quite well but its not their current focus on creating or allowing others to create small finetuned and optimized models.

In the race they are in, its still highly beneficial to be the frontier model.

I use the frontier model every single day through my company and my company happily spends these tokens.

It makes a huge difference if different people can use one interface to do everything.

Big models give you fundamental things: A lot of facts/context, usability (you might be a native english speaker, don't underestimate how hard it is for A LOT of people to formulate what they want/need in english only AND a low complexity.

No one needs to build a router and x sub models and a router architecture. You literaly just have an API, you might choose the model and the effort but thats it.

I find this current state of the art a LOT more telling on the current progress we are in than anything else. I'm confident that small optimized models will become a lot more relevant like lets say java + english + a second language + spring boot + postgresql. It might also be beneficial in the long term to finetune with your project details.

For a lot of very technical non human interfacing things, finetuning is happening left and right.

But what i find very interesting is emerging complexity capabillity. I believe that fables skill to hold more topics and combine more complex solutions together is because of its parameter size.

We will have to figure out if we can extract this complexity out of it while reducing the training data in a way that the training data focus more on thinking. Plenty of smaller thinking models show that this is doable.

Btw. Mixture of Experts is for sure not optimal for this, but it already is a form of optimized sub models. Perhaps we might just have MoE with a million experts in the future. One per lanuage + area of expertise etc.

Also don't forget: IF AGI is coming through a current frontier model, you will let it work for hours, days and weeks on one problem completly independent of any human input and it will be better than a human. If they reach this before a collapse, we are done and they 'won'. For this you need big frontier models.