Hacker News new | ask | show | jobs
by dgritsko 2 days ago
This isn't quantized, right? Just a smaller context?
2 comments

Its 256k context window. Quantization is orthogonal. We cant really tell directly so it could be quantized.
The model is already natively MXFP4-quantized during training, so there is no quality loss.