Latency can be just as important as overall throughput, especially for inference providers like Groq and Cerebras.