Hacker News new | ask | show | jobs
by KolibriFly 15 days ago
Another important thing to keep in mind here is the software. Nvidia has been optimizing their software stack for decades to squeeze every clock cycle out of the hardware under any limits. TensorRT alone does straight up magic. Apple is just starting out with MLX, and their hardware is often idling not because it is worse per watt but because the compiler does not know how to optimally load the ALU blocks for specific graphs yet