| HN Mirror

Y	Hacker News new \| ask \| show \| jobs


	by PoignardAzur 299 days ago
	Does Gemma use any specific scheme to compress embeddings? Which have you considered? For instance, it's well-known that transformer embeddings tend to form clusters. Have you considered splitting the embedding table into "cluster centroid" and "offset from centroid" tables, where the later would presumably have a smaller range and precision?