Hacker News new | ask | show | jobs
by om8 9 days ago
This project needs webgpu -- I did it on cpu about a year ago.

My demo uses 2 bit quantization to run llama3 models on any device with enough ram.

https://galqiwi.github.io/aqlm-rs/