|
|
|
|
|
by kingstnap
25 days ago
|
|
A consumer CPU like a 285k caps out around 130 GB/s of memory bandwidth. Each of its 24 cores can do two 8 wide FMA ops per cycle. Lets say holding a continuous 4 GHz clock speed. This works out to over 1.5 trillion 32 bit floating point multiplies per cycle. If you are doing vector matrix multiplies (like in single token no batching). `xW` then each weight loaded sort of gets used in 1 multiplication and 1 addition. Doing the math you can clearly see even if each weight were just 1 byte you can at most load 130 billion of them in a second from memory. But in the same timespan you could have done over 1.5 trillion multiplications. So you are still memory bound. |
|