|
|
|
|
|
by in-silico
121 days ago
|
|
In theory you do lose information compared to parameters with more bits. In practice, neural networks aren't able to store much more than 2-4 bits of useful information per parameter (regardless of the precision), so models like this are mostly getting rid of redundancy. |
|