Hacker News new | ask | show | jobs
by angusturner 4 hours ago
I find the presentation of this as "factoring training" very confusing.

Factoring is something we do to distributions, which we then parameterize with our neural net.

An autoregressive model pertains to a certain factorisation of a joint distribution. And if we choose this factorisation it may have consequences for our training and sampling etc.

But to say "most models on factor generation".. I dont quite get this.

This would be better work if it didnt claim it was a new pretraining axis also. Its another way of scaling compute right?

If it were a new axis, IWAE would have an equal claim to it (as others have pointed out).