Hacker News new | ask | show | jobs
by kingstnap 25 days ago
A consumer CPU like a 285k caps out around 130 GB/s of memory bandwidth.

Each of its 24 cores can do two 8 wide FMA ops per cycle. Lets say holding a continuous 4 GHz clock speed.

This works out to over 1.5 trillion 32 bit floating point multiplies per cycle.

If you are doing vector matrix multiplies (like in single token no batching). `xW` then each weight loaded sort of gets used in 1 multiplication and 1 addition.

Doing the math you can clearly see even if each weight were just 1 byte you can at most load 130 billion of them in a second from memory.

But in the same timespan you could have done over 1.5 trillion multiplications.

So you are still memory bound.