Hacker News new | ask | show | jobs
by mandelken 9 days ago
On Strix Halo, have a look at the work antirez is doing with dwarfstar (antirez/ds4) for deepseek v4 flash at reasonable speeds and quality.
1 comments

I've tried it, but, it's not reasonable speed, at all. It's 9-13 tokens per second, which is not usable interactively and not worth using for long-running API stuff when DeepSeek V4 Pro is so cheap via their API.

Laguna S 2.1 runs at 15-28 tokens per second, depending on context and...something about how long it's been running, which is very comfortable for chatting, but still not usable for interactive agentic coding. Their `pool` agent just times out when I try to use it with the Strix Halo-hosted instance of the model.