|
|
|
|
|
by mercutio2
18 days ago
|
|
For agentic work, you just cache the prefill kv cache of the relevant system prompts. TTFT is a little slow (10s of seconds, oh no!) the first time you boot up a new harness. I keep hearing people make this claim that TTFT is a problem, and… it just isn’t, if you’re running oMLX. Folks in my camp keep saying this, and folks in your camp keep beating a drum we tell you isn’t resonating. Not sure why I keep bothering to argue; you can’t buy a high RAM Mac Studio like mine anymore. |
|
They're not usable for deployment. They're perfectly fine for "enthusiast" low-end usage with 10-30B models, but the same goes for almost every dGPU made in the last 10 years. Your Mac Studio cannot run frontier LLMs at an interactive speed, even Apple has given up on using it as an inference backbone.