Hacker News new | ask | show | jobs
by in-silico 121 days ago
In theory you do lose information compared to parameters with more bits.

In practice, neural networks aren't able to store much more than 2-4 bits of useful information per parameter (regardless of the precision), so models like this are mostly getting rid of redundancy.