Y
Hacker News
new
|
ask
|
show
|
jobs
by
Iolaum
2 days ago
As long as they are transparent about what quant they serve the model and any other optimization they do that also affects performance of inferred tokens.