Hacker News new | ask | show | jobs
by jazzpush2 77 days ago
A Codon-based model is cool. I know NVIDIA is building quite a large one.

At GTC they showed an SAE they built on a smaller version of it, allowing you to see what their model learned: https://research.nvidia.com/labs/dbr/blog/sae/