Hacker News new | ask | show | jobs
by Iolaum 2 days ago
As long as they are transparent about what quant they serve the model and any other optimization they do that also affects performance of inferred tokens.