|
|
|
|
|
by solarkraft
63 days ago
|
|
I used to be so bullish on Cerebras, being pretty certain that specialized chips would eventually dethrone NVidia for inference hardware. Multiple years have passed since then. GPT spark got me excited again, but somehow that seems to have faded right back into obscurity. Can somebody explain why there’s so little apparent progress here despite the theoretically massive advantage? Can I still expect this to happen eventually? |
|
In the future, they plan hybrid implementations, to be able to serve large models better, e.g.
"AWS. We signed a binding term sheet with Amazon Web Services for AWS to become the first hyperscaler to deploy Cerebras systems in its data centers. Deployment in AWS data centers will require us to meet strict standards for performance, scale, and reliability.Pursuant to the term sheet, we will create a co-designed, disaggregated inference-serving solution that will integrate AWS Trainium3 chips with Cerebras CS-3 systems, connected via high-bandwidth networking, to partition inference workloads across Trainium3 and CS-3. Each system will perform the type of computation at which it most excels. The approach is expected to deliver 5 times more token throughput in the same hardware footprint, at up to 15 times faster speeds compared to leading GPU-based solutions as benchmarked on leading open-source models."