Hacker News new | ask | show | jobs
by Kirby64 11 days ago
Less impressive than you might think here. I can't imagine how small this model must be if it only has 1.46M cells. Maybe 1M? Probably less. Sure the token/s decode is cool, but this size isn't capable of doing anything. For context: this is a 45nm process.

In comparison, Taalas' implementation (which everyone likes to talk about) is in 6nm, is 815mm^2, and only serves a 8B parameter ~4-bit quant model.

So, 815mm^2 in 6nm is roughly equivalent to ~6112mm^2 in 45nm. If we assume everything scales exactly the same, 4mm^2 would be ~1500x smaller. 1500x smaller means at best we're talking about a ~5M parameter model. I don't know how you'd get a 5M param model (with multiple bits) in 1.46M transistors.