There is a lot of research out there with answers to your questions. GPU supply chains are extremely constrained, making it a surprisingly tractable problem at least for now.
How do you tell the difference between inference and training workloads from outside? The same GPUs can be used for both and the pacing suggested here relates to training and not inference as far as I can tell.