| HN Mirror

Y	Hacker News new \| ask \| show \| jobs


	by modeless 248 days ago
	The models (weights and activations and caches) can fill all the memory you have and more, and to a first (very rough) approximation every byte needs to be accessed for each token generated. You can see how that would add up. I highly recommend Andrej Karpathy's videos if you want to learn details.

2 comments

pfortuny 248 days ago

A very simplified version is: you need all the matrix to compute a matrix x vector operation, even if the vector is mostly zeroes. Edit: obviously my simplification is wrong but if you add up compression, etc… you get an idea.

link

rs186 248 days ago

Would you mind specifying which video(s)? He has quite a lot of content to consume.

link