|
|
|
|
|
by ryao
581 days ago
|
|
This is on llama 3.1 405B. Inferencing is memory bandwidth bound. Add more GPUs on a batch size 1 inference problem and watch it run no faster than the memory bandwidth of a single GPU. It does not scale across the number of GPUs. If it could, you would see clusters of Nvidia hardware outperforming Cerebras’ hardware. That is currently a fantasy. |
|
[1]: https://lmsys.org/blog/2024-07-25-sglang-llama3/?ref=blog.ru...
[2]: https://www.snowflake.com/engineering-blog/optimize-llms-wit...