Hacker News new | ask | show | jobs
by Amekedl 14 days ago
No.

1. "Grokking" was shown on 4-digit modular arithmetic with a 1-layer transformer; this article extrapolates it to AGI and a $10B training run with exactly zero intermediate evidence.

2. The "small dataset" is 25 trillion tokens - literally the size of current frontier training sets - but calling it small sounds revolutionary.

3. BabyLM has spent 4 years failing to produce grokking on constrained data; the paper gets a footnote saying "those models were too small," which is unfalsifiable until someone burns $10B.

4. Chain-of-thought is already empirically required for frontier performance - it's expensive, bizarre, and nobody predicted it - yet somehow we're supposed to bet the farm on a phenomenon that has never scaled past arithmetic. We need that data, even if it is just "Actually, ..."

5. If you want to chase "recurrent depth", loop transformers rumored in Mythos/Fable are at least grounded in actual engineering; grokking-at-scale is just vibes ai bro science.

More data is and will always be the answer. Why are all labs distilling from each other?

1 comments

FYI, AI-written comments are banned on Hacker News: https://news.ycombinator.com/newsguidelines.html

> Don't post generated text or AI-edited text. HN is for conversation between humans.

? How could Grokking be made workable? Mechanistic Interpretability, like using SAEs (EleutherAI did), also methods Heretic employs to decensor, surely you can see some "modding" does indeed achieve things.

FYI, I prefer discussion over being mad