Hacker News new | ask | show | jobs
by CamperBob2 43 days ago
Not OP, but I am seeing up to 260 tokens/second output at c=1 with the recipe at https://github.com/local-inference-lab/rtx6kpro/blob/master/... using 4x 6k cards. Average is more like 200.

There may be a way to get the 2-bit quantized version running even faster on a pair of them.

1 comments

Thank you. Useful to know. Clipped on top by reduce, I assume.
I think so. The machine I'm using runs at Gen4 x8, while the cards can take advantage of Gen5 x16.