Y
Hacker News
new
|
ask
|
show
|
jobs
by
te0006
43 days ago
Interesting setup. What GPU(s)/VRAM, CPU and RAM are you using for the 122B model, with which quantization, and what token rates do you achieve for prefill and generation?