Hacker News new | ask | show | jobs
by alienbaby 25 days ago
I'm interested if anyone knows how much legwork the assumed 60% cache hit, plus running a quantised model is doing? Esp. compared to what the headline half implies is a full fat GLM5.2