Hacker News new | ask | show | jobs
by jaxytee 23 days ago
Running large MOE models won't saturate the EPYC memory.

You'll never see that 200 Gb/s running inference.

1 comments

Because it'll bottleneck at the CPU? Or because 200 bB/s is overkill when running a model (that fits in memory)?