|
|
|
|
|
by cmrdporcupine
21 days ago
|
|
Until RAM prices drop and can economically get machines with 256GB, 512GB and higher bandwidth... I frankly think the local AI story is going to be still fairly muted for most people. My Spark can do Qwen3.6 MoE A3B at 60 to 70-ish token/second and that's really good, but there's limits the usefulness of that model. It's not useful for coding, in any case. Once people can run something like GLM 5.2 at lower quants (512GB could do a passable job), then I think the story changes. Whether we ever see DRAM as cheap as it was ever again, I don't know. |
|
That doesn't mean that local models are useless though! If Mythos/Sol is an ASI that threatens to take your job and turn you into paperclips, then Qwen/Gemma is an old-fashioned office secretary that loyally helps you with tasks but doesn't have a good grasp of details. Every white-collar worker 50 years ago would have killed to have a hard-working personal secretary.