Hacker News new | ask | show | jobs
by kees99 9 days ago
Qwen is quite slow in general.

I'm getting ~0.5 tps from [qwen], and ~10 tps from [gemma] - roughly the same size, same quant, same hardware (8-core CPU) same software (llama.cpp).

[qwen] Qwen3.6-27B-Q4_K_M.gguf

[gemma] gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf

1 comments

Quite different models though Qwen 3.6 model has 27B active parameters. gemma-4 has 26B parameters with 4B active at any given time. You should compare with Qwen 3.6 35B A3B