Hacker News new | ask | show | jobs
by WhitneyLand 122 days ago
A bit misleading to say they take 14x less memory, no one is doing inference with 16-bit models.