Hacker News new | ask | show | jobs
by ejpir 22 days ago
this is not true. q2 Deepseek flash works on AMD Strix halo with pretty good results. Benchmark except:

ctx_tokens,prefill_tokens,prefill_tps,gen_tokens,gen_tps,kvcache_bytes 2048,2048,202.02,128,15.31,52184460 4096,2048,211.03,128,14.64,80373132 6144,2048,208.04,128,14.59,108561804 8192,2048,200.78,128,14.43,136750476 10240,2048,203.04,128,14.37,164939148 12288,2048,200.82,128,14.27,193127820 14336,2048,198.62,128,14.22,221316492 16384,2048,196.14,128,14.20,249505164 18432,2048,189.48,128,14.13,277693836 20480,2048,186.59,128,14.06,305882508 22528,2048,183.88,128,13.99,334071180 24576,2048,183.38,128,13.92,362259852 26624,2048,181.57,128,13.87,390448524 28672,2048,183.46,128,13.80,418637196 30720,2048,181.80,128,13.73,446825868 32768,2048,175.93,128,13.55,475014540 34816,2048,175.42,128,13.46,503203212

https://kyuz0.github.io/strix-halo-ds4-toolbox/

1 comments

I said, "Like, DS4 (a 1-bit quantization of DeepSeek V4 Flash) runs at something like 9-13 tokens/second, with a loooong time to first token."

Which almost exactly matches the benchmark you just linked. Looks like it's possible to goose it to 15 tokens per second with a tiny context, but why would I want a giant model with a 2k context? DeepSeek is too big to be fast enough on a Strix Halo.

To be clear, if you think that's comfortable for interactive use, more power to you. But, I'm not waiting for that. I'll pay DeepSeek to host it for me. Their token prices are quite cheap and their cached tokens are even cheaper...and they have the most effective caching in the business, as far as I can tell. Even naively using the API you get 80-90% cached token rate. If you use Reasonix, you get ~98% cached token rate. I just built a feature for an app I'm working on for $0.10 for 20 minutes of work. Not bad at all.