Are local models more snappy? I'm at a stage where I can work with the output of LLMs, and the next win is really just getting things written out quickly.
Depends on your setup. If you drop $10k on an RTX Pro 6000 then yeah Qwen 35B MoE will absolutely fly.
If you have a pair of 3090s and run Qwen 27B, or an old Threadripper with heaps of system RAM and Deepseek or MiniMax or Kimi, no it won't be as fast as Claude.
Most local LLM nerds are not running locally for superior speed, we're doing it for sovereignty and/or privacy, or maybe just because it's fun which accidentally became useful this year.
If you have a pair of 3090s and run Qwen 27B, or an old Threadripper with heaps of system RAM and Deepseek or MiniMax or Kimi, no it won't be as fast as Claude.
Most local LLM nerds are not running locally for superior speed, we're doing it for sovereignty and/or privacy, or maybe just because it's fun which accidentally became useful this year.