Hacker News new | ask | show | jobs
by _boffin_ 33 days ago
That statement gives a lot of insights into the possible model size. Llama 3.1 405b runs @ ~900t/s: https://www.cerebras.ai/blog/llama-405b-inference

From initial vibe research (which is totally not correct by any means) ~13.5k concurrent streaming clients capacity