Hacker News new | ask | show | jobs
by gHA5 16 hours ago
I always wonder how often people in charge of massive compute tasks like LLM pre-training have messed up some detail that invalidates or fails to persist the results and only realise afterwards.