|
|
|
|
|
by ux266478
1 day ago
|
|
It's not that surprising to me. Most of the innovation in Chinese models has been in efficiency gains and optimizations. K3 coming from the factory in MXFP4 weights is a pretty relevant factor. Big performance gap probably also due to Moonshot doing QAT. Throw in the fact that Musk has easier access to compute, and I think you have your answer on the disparity. |
|
When OAI released gpt-oss it was released as an mxfp4 checkpoint.
OAI, Ant, et al are also obviously employing QAT.