These token numbers look off for a 5090. For comparison, an rtx 5000 Blackwell SFF with just 470gb/s bandwidth gets me 30 tok/s on gemma4-31B-QAT. Almost 60 tok/s with MTP enabled. A 5090 should get you much more than that!
Thanks for flagging, this is on my local rig and it's driving my display too. I'm curious now, will take a closer look. These are the tok/s as reported by LMStudio.
EDIT: Updating llama.ccp gets me 58 tok/s on Gemma 31b
The Qwen-35B-A3B numbers are even weirder, did you drop a digit? I get half of that speed on a AMD Radeon RX 5500 XT (RADV NAVI14) (8192 MiB) (unsloth/Qwen3.6-35B-A3B-MTP-GGUF:Q6_K with q8_0 context).
bahaha, LMStudio's badge looks a lot like a B but the tooltip says it is indeed an 8 (I just assumed it was B for "byte", but that's a poor initialism considering "bit").
EDIT: Updating llama.ccp gets me 58 tok/s on Gemma 31b