|
|
|
|
|
by nok22kon
28 days ago
|
|
the scaling laws work within a "generation". but what about across them? GPT-3 was 175B, models like Gemma4 with 31B vastly outperform it, so there is more to it as Karpathy noted, the initial GPTs were trained on complete garbage (literally, the average document from the Common Crawl is random nonsense), yet they worked. now we can use present LLMs to curate the data for the next generation |
|