Hacker News new | ask | show | jobs
by basiccalendar74 1 day ago
Weights will be fetched 4-bit from memory, but compute happens in FP8 (W4A8) or BF16 (W4A16). This is still better than fetching 8-bit weights from memory.

Both W4A8 and W4A16 schemes are supported by Hopper GPUs and commonly used to serve mxfp4 Kimi models on Hopper.