Hacker News new | ask | show | jobs
by broabprobe 17 days ago
I run the same setup Gemma 4 26B on a 2013 Mac Pro (dual graphics cards but they're useless for this). I also get about 5 t/s. It's perfectly serviceable for some tasks!
2 comments

I bought a trashcan Mac Pro on a whim last week ($120 in eBay!) and did some reading about them—it turns out people recently started using the GPUs to run models @ 20-30 tok/s.

I'm excited to get my mitts on it on Friday when it finally arrives.

Here's some of the resources I came across if you're interested in reading.

https://echalupa.com/blog/mac-pro-6-1-llama-cpp-firepro-d300...

https://matthewgribben.com/blog/mac-pro-6-1-llama-cpp-firepr...

oh very exciting! Thanks for sharing these sjs382! A bummer the models can't be run on both the GPUs and CPU so you could run much larger models. A year ago I was running gpt120b as it easily fit in 128gb ram. But now I'm only running gemma 26b and wishing I had stuck with 64gb ram since 128gb gets throttled. Not regretting buying 128gb 1.5 years ago though!
What is it useful for at slow speed?