Hacker News new | ask | show | jobs
by te0006 43 days ago
Interesting setup. What GPU(s)/VRAM, CPU and RAM are you using for the 122B model, with which quantization, and what token rates do you achieve for prefill and generation?