|
|
|
|
|
by Kirby64
11 days ago
|
|
Less impressive than you might think here. I can't imagine how small this model must be if it only has 1.46M cells. Maybe 1M? Probably less. Sure the token/s decode is cool, but this size isn't capable of doing anything. For context: this is a 45nm process. In comparison, Taalas' implementation (which everyone likes to talk about) is in 6nm, is 815mm^2, and only serves a 8B parameter ~4-bit quant model. So, 815mm^2 in 6nm is roughly equivalent to ~6112mm^2 in 45nm. If we assume everything scales exactly the same, 4mm^2 would be ~1500x smaller. 1500x smaller means at best we're talking about a ~5M parameter model. I don't know how you'd get a 5M param model (with multiple bits) in 1.46M transistors. |
|