Hacker News new | ask | show | jobs
by l332mn 46 days ago
Mind sharing your setup? I also have dual 3090s, but getting nowhere close to 300k context limits with 4 bit quantized models at that size (using vllm).