| HN Mirror

Y	Hacker News new \| ask \| show \| jobs


	by leakyfilter 377 days ago
	Raw gemm computation was never the real bottleneck, especially on the newer GPUs. Feeding the matmuls i.e memory bandwidth is where it’s at, especially in the newer GPUs.