|
|
|
|
|
by cthalupa
25 days ago
|
|
I’ve got a 128gb m5 max mbp and two sparks. For my real-world use cases, a single spark running DS4 Flash will have fully responded by the time my Mac has even started generating tokens. I figured I’d have more generation heavy work when I also bought the Mac, but it has done very little work running LLMs since I got the first Spark. I mostly run them clustered for DS4 and am quite happy with the performance, and the cost isn’t that much more for two than the MBP while giving me double the unified memory. I’ll probably pick up a third to run multiple smaller models. I don’t understand why people would buy a halo over a spark at comparable prices, particularly because if you want to cluster, the cx7 be beats the shit out of them when it comes to latency and throughput |
|