|
|
|
|
|
by b33j0r
1195 days ago
|
|
It seems pretty transparent, as I think you might be implying in part, that attempting to leap-frog without directly copying training data or model weights is a temporary optimization for “bootstrapping” teams. I see this pretty directly in stablelm releasing both a base model, and a tuned model… which is not based on the base ;) There is a goldrush to get training sets worth using, and if something 90% quality gets your models on the map quickly, it’s an attractive option. As they say, attention is all you need. Training in the 7B range is a lot cheaper than I expected. Fine-tuning almost negligible—if you have clean data, which has always been the expensive part. Most humans expect more than $0.000003 per token as compensation for _your_ dataset collection. |
|