Hacker News new | ask | show | jobs
Re-quantizing a local LLM 14x faster by skipping the tensors that didn't change (andreaborio.substack.com)
8 points by andreaborio 49 days ago