|
this is not true. q2 Deepseek flash works on AMD Strix halo with pretty good results. Benchmark except: ctx_tokens,prefill_tokens,prefill_tps,gen_tokens,gen_tps,kvcache_bytes
2048,2048,202.02,128,15.31,52184460
4096,2048,211.03,128,14.64,80373132
6144,2048,208.04,128,14.59,108561804
8192,2048,200.78,128,14.43,136750476
10240,2048,203.04,128,14.37,164939148
12288,2048,200.82,128,14.27,193127820
14336,2048,198.62,128,14.22,221316492
16384,2048,196.14,128,14.20,249505164
18432,2048,189.48,128,14.13,277693836
20480,2048,186.59,128,14.06,305882508
22528,2048,183.88,128,13.99,334071180
24576,2048,183.38,128,13.92,362259852
26624,2048,181.57,128,13.87,390448524
28672,2048,183.46,128,13.80,418637196
30720,2048,181.80,128,13.73,446825868
32768,2048,175.93,128,13.55,475014540
34816,2048,175.42,128,13.46,503203212 https://kyuz0.github.io/strix-halo-ds4-toolbox/ |
Which almost exactly matches the benchmark you just linked. Looks like it's possible to goose it to 15 tokens per second with a tiny context, but why would I want a giant model with a 2k context? DeepSeek is too big to be fast enough on a Strix Halo.
To be clear, if you think that's comfortable for interactive use, more power to you. But, I'm not waiting for that. I'll pay DeepSeek to host it for me. Their token prices are quite cheap and their cached tokens are even cheaper...and they have the most effective caching in the business, as far as I can tell. Even naively using the API you get 80-90% cached token rate. If you use Reasonix, you get ~98% cached token rate. I just built a feature for an app I'm working on for $0.10 for 20 minutes of work. Not bad at all.