|
|
|
|
|
by amelius
1 day ago
|
|
> Without training cost you can infer only the marginal cost of serving this kind of models. Which is by far the most interesting number of the two. > Moreover, you don't know the actual size of closed models (what if Fable is a 10T model? What if it's 1T?) If you get close in output quality, then does that matter? |
|
When you're trying to estimate/infer the costs of serving the tokens and even include the cost of training the weights in order to output tokens then yeah, why wouldn't that matter?