| HN Mirror

Y	Hacker News new \| ask \| show \| jobs


	by jwan584 1043 days ago
	A helpful paper with the full recipe Cerebras uses to train LLMs and their process including: - Extensively deduplicated dataset (SlimPajama) - Hyperparameter search using muP - Variable sequence length training + ALiBi - Aggressive LR decay