Y
Hacker News
new
|
ask
|
show
|
jobs
by
jaxytee
23 days ago
Running large MOE models won't saturate the EPYC memory.
You'll never see that 200 Gb/s running inference.
1 comments
edg5000
23 days ago
Because it'll bottleneck at the CPU? Or because 200 bB/s is overkill when running a model (that fits in memory)?
link