Hacker News new | ask | show | jobs
by johndough 4 days ago
Why not? The model is 2.8T parameters with native MXFP4, which is 1400GB.
1 comments

I would figure it needs a lot of kV cache and context size which can be much bigger than the model itself