| HN Mirror

Y	Hacker News new \| ask \| show \| jobs


	by Taek 1137 days ago
	75B tokens is not really enough data to make an intelligent model. Llama was trained on over 1 trillion tokens. And yes, training on top of LLaMA could introduce a lot of unexpected behavior, but that's just where the State-of-the-Art is today