Hacker News new | ask | show | jobs
by b33j0r 1195 days ago
It seems pretty transparent, as I think you might be implying in part, that attempting to leap-frog without directly copying training data or model weights is a temporary optimization for “bootstrapping” teams.

I see this pretty directly in stablelm releasing both a base model, and a tuned model… which is not based on the base ;)

There is a goldrush to get training sets worth using, and if something 90% quality gets your models on the map quickly, it’s an attractive option. As they say, attention is all you need.

Training in the 7B range is a lot cheaper than I expected. Fine-tuning almost negligible—if you have clean data, which has always been the expensive part.

Most humans expect more than $0.000003 per token as compensation for _your_ dataset collection.