|
|
|
|
|
by pornel
19 hours ago
|
|
IANAMLE, but there is "grokking" that makes models learn to actually generalize, even after you give them enough parameters that would let them memorize the dataset: https://en.wikipedia.org/wiki/Grokking_(machine_learning) High-dimensional gradient descent behaves very differently than the simplified 3d visualisations we use to demonstrate it, and has lots of ways out of local minima: https://www.youtube.com/watch?v=NrO20Jb-hy0 so it seems like there is a benefit to giving models more space to learn in rather than forcing them to compress the knowledge from the start. |
|