My demo uses 2 bit quantization to run llama3 models on any device with enough ram.
https://galqiwi.github.io/aqlm-rs/