Hacker News new | ask | show | jobs
by reissbaker 1 day ago
No, OpenAI did not publish how they trained o1, and at the time there was significant misunderstanding and belief in the research community that they were using some kind of Monte-Carlo tree search. DeepSeek figured out GRPO on their own. Similarly, while others invented MoEs, DeepSeek's ultra-sparse variants were extremely novel, to the point where the revelation of how efficient they were to train temporarily collapsed Nvidia's stock.

Regardless I think it's impossible to believe that most LLM research was done by closed labs that don't publish, especially Anthropic (who missed out on and copied two of the largest pieces of important research of the last several years), and that none was done by open labs like DeepSeek, and that the open labs are just copycats. It's quite clear that isn't the case.