Hacker News new | ask | show | jobs
by pornel 19 hours ago
IANAMLE, but there is "grokking" that makes models learn to actually generalize, even after you give them enough parameters that would let them memorize the dataset:

https://en.wikipedia.org/wiki/Grokking_(machine_learning)

High-dimensional gradient descent behaves very differently than the simplified 3d visualisations we use to demonstrate it, and has lots of ways out of local minima:

https://www.youtube.com/watch?v=NrO20Jb-hy0

so it seems like there is a benefit to giving models more space to learn in rather than forcing them to compress the knowledge from the start.

1 comments

Exactly, I was going to suggest the Welch Labs videos on grokking. Especially the newer one at https://youtu.be/D8GOeCFFby4?si=yLI9zzcjsnEELUqy . They are really well done and really eye-opening.