|
|
|
|
|
by londons_explore
19 days ago
|
|
There are various difficulties with renting GPUs, especially if your setup is very custom. The competitor would have to port their training systems to your specific network architecture, system design, rdma Vs ethernet vs infiniband Vs nvlink etc. Getting it running might not be too hard, but getting it running efficiently and making good use of all those flops will require considerable human effort and wall time. Add that to the fact most frontier labs seem to have a single huge training run - and to my knowledge nobody has figured out how to distribute that training run between data centers effectively. |
|