|
|
|
|
|
by chatmasta
11 days ago
|
|
Inevitably the CSPs will make their own hardware, especially as we start to see specialized chips for specific models or generic inference. This is already happening with Google and TPUs. It’s easier for the CSPs to move into hardware than it is for Nvidia to move into cloud hosting. Although as a middle ground I’ve been quite happy with Nvidia Brev for on-demand GPU instances from a select marketplace of CSP offerings. It’s a well kept secret IMO — great product (from an acquisition iirc). |
|
Also, not sure how well CSPs inference stack is compared with vllm + nvidia. A lot of open weight models uses MoE, making the inference stack more complex.