|
|
|
|
|
by hedgehog
7 days ago
|
|
Yes that's what I've read. As far as I know the approach should transfer well to hybrid model architectures like modern Qwen and sizes like 27B by using multiple chips. LoRA-steered Qwen 27B at 10K+ tokens per second would be transformative for some workflows. |
|