|
|
|
|
|
by aaulia
16 days ago
|
|
I tried Qwen MoE a while back. Using my 8GB RX470, somehow got 10+ token/sec, lot's of trial and error with llama.cpp config, and it's still slow to be used for my usecase. Even at 12 to 16 IMO it's slow. For chat, maybe it's enough, but for any other tasks, it's not viable IMO |
|