Y
Hacker News
new
|
ask
|
show
|
jobs
by
MarsIronPI
122 days ago
It's because they're natively trained with 1 bit, so it's not losing anything. Now, the question might be how they manage to get decent predictive performance with such little precision. That I don't know.
1 comments
syntaxpr
122 days ago
Not training. Transposing rows/columns of matrices to group 128 parameters with similar (shared) scale factor. Qwen-3 model.
link
MarsIronPI
122 days ago
I'm not sure what you mean. Could you please elaborate?
link