| HN Mirror

Y	Hacker News new \| ask \| show \| jobs


	by ryao 581 days ago
	This is on llama 3.1 405B. Inferencing is memory bandwidth bound. Add more GPUs on a batch size 1 inference problem and watch it run no faster than the memory bandwidth of a single GPU. It does not scale across the number of GPUs. If it could, you would see clusters of Nvidia hardware outperforming Cerebras’ hardware. That is currently a fantasy.

1 comments

This two sources[1][2] shows 1500-2500 token/per second on 8*H100.