Y
Hacker News
new
|
ask
|
show
|
jobs
by
arjie
43 days ago
Vouched your comment. Very cool. What are you running on to get 190 tok/s? I get 400 tok/s at c=4 but c=1 is slower than you.
2 comments
CamperBob2
43 days ago
Not OP, but I am seeing up to 260 tokens/second output at c=1 with the recipe at
https://github.com/local-inference-lab/rtx6kpro/blob/master/...
using 4x 6k cards. Average is more like 200.
There may be a way to get the 2-bit quantized version running even faster on a pair of them.
link
arjie
43 days ago
Thank you. Useful to know. Clipped on top by reduce, I assume.
link
CamperBob2
42 days ago
I think so. The machine I'm using runs at Gen4 x8, while the cards can take advantage of Gen5 x16.
link
mtone
43 days ago
I am using the `voipmonitor/vllm:lucifer` docker from the RTX6K discord community discussed at the same link the other commenter posted. It is based around this PR
https://github.com/vllm-project/vllm/pull/43477
link
arjie
43 days ago
Ah I’m on the same PR just behind. Thank you.
link
There may be a way to get the 2-bit quantized version running even faster on a pair of them.