Hacker News new | ask | show | jobs
by ChanderG 1 day ago
Why? Why is the premise that Fine-tuned models should be geared towards new discoveries?

The point of Fine-tuning small models is for specific downstream tasks, which SOTA models can do, but at higher costs. It is purely an economic play, not an attempt at pushing boundaries of SOTA.

2 comments

I'm not suggesting that fine-tuned models don't have their place, all I'm saying is that the constant drumbeat of "cheap model X beats more expensive model Z" completely misses that the more expensive model is capable of doing more things at a higher level.

If the appropriate qualifiers were added to say "cheap model X does better at test Y than expensive model Z when we fine tune X to take Y test of existing knowledge" then it would be a more accurate statement, but naturally less impressive.

Maybe because people who are target audience don’t need to have it spelled out like that?

People who are not really into it, don’t care.

> I'm not suggesting that fine-tuned models don't have their place, all I'm saying is that the constant drumbeat of "cheap model X beats more expensive model Z" completely misses that the more expensive model is capable of doing more things at a higher level.

What if you have to do the task a billion times? Which model will you choose?

Speed play also you can get much faster responses with 9b model.