Hacker News new | ask | show | jobs
by reissbaker 1 day ago
I think it's pretty hard to hold that worldview: Anthropic couldn't ship a reasoning model until they copied DeepSeek R1's homework, and they've all copied DS-style super-sparse MoEs at this point too.
2 comments

That’s a really good point. Folks really need to read the papers coming out of these Chinese labs. Every paper from the DeepSeek team has been a step change.
With slightly different cherry-picking, you could equally well claim that DeepSeek couldn't ship a reasoning model until they copied the idea from OpenAI's o1-preview, and they also copied MoEs from Google Brain/Jagellonian University https://arxiv.org/abs/1701.06538 way back in 2017, too!

But ultimately these were ideas floating around in the air, if one group hadn't done the experiment, someone else would have.

No, OpenAI did not publish how they trained o1, and at the time there was significant misunderstanding and belief in the research community that they were using some kind of Monte-Carlo tree search. DeepSeek figured out GRPO on their own. Similarly, while others invented MoEs, DeepSeek's ultra-sparse variants were extremely novel, to the point where the revelation of how efficient they were to train temporarily collapsed Nvidia's stock.

Regardless I think it's impossible to believe that most LLM research was done by closed labs that don't publish, especially Anthropic (who missed out on and copied two of the largest pieces of important research of the last several years), and that none was done by open labs like DeepSeek, and that the open labs are just copycats. It's quite clear that isn't the case.