Hacker News new | ask | show | jobs
by drob518 123 days ago
I’m really curious how this scales up. Bonsai delivers an 8B model in 1.15 GB. How large would a 27B or 35B model be? Would it still retain the accuracy of those large models? If the scaling holds, we could see 100+B models in 64 GB of RAM.
1 comments

Also depends on how expensive training these models is. It's probably at least as expensive as full precision models, otherwise they would have mentioned it.
My guess is the training process is their secret sauce...
Yes, but their training speed is not secret. If their process were fast, they would have said so.