Hacker News new | ask | show | jobs
by bahmboo 18 days ago
To be fair it's "only" half the throughput of a 4090 and a third of an RTX 6000. Significant but not an order of magnitude.
3 comments

Those are the ratios for memory bandwidth, but the GPUs have a much higher ratio for compute, and that affects prefill rate / TTFT, right?
For local inference, the difference between 25t/s and 70t/s is a lot. For some models I struggle to even reach 15t/s. And "some models" aren't even large models, Gemma 4 13b has this issue for some reason. For stuff like Qwen3.6-27B I can hardly reach 10t/s, even with fully custom inference made by Fable 5!
An old ada Rtx 6000 maybe. A Blackwell RTX Pro 6000 is an order of magnitude faster and has 96gb.
That's not what I'm seeing. It is much faster but not an order of magnitude. Not trying to be pedantic, only setting expectations.

"The Blackwell RTX PRO 6000 provides up to 1,792 GB/s of memory bandwidth, while the 40-core Apple M5 Max tops out at 614 GB/s"

Sorry, thought we were talking about tokens. M5 Max is great for bandwidth and I’m looking forward to seeing what Apple does for AI inference in the M7. The 6000 kills everything else when it comes to TTFT and tokens/s.
For sure. Clearly Nvidia mops the floor with the competition. I'm looking forward to M6/M7 and to see if Apple wants a bigger piece of the pie.