| This kind of systematic distillation by a competitor can allow them to fast-follow you and pick up capabilities. If you've invested in expensive capabilities training, of course you don't want this, so it's in Anthropic's economic interest to hinder it however they can, and that's enough to explain their behaviour here. Anthropic seems to genuinely care about safety though, which for the rest of us means not having models that enabling easier cyberattacks, targeted scams, and the rarer but more severe risks like people trying to create and release new pathogens. This means walking a tight line, especially as models become more capable, and often wrapping a model in layers of defences against misuse. If those capabilities transfer to a closed competitor model, all bets are off in terms of whether the competitor will apply the same defences. If those capabilities transfer to an open weight model, not only will there be no ring of defences around the model, any defences you put into the model itself can easily be stripped away. So although it's nice to have capable open models, it will increasingly bad for us all if open models keep fast-following closed model capabilities as they have been, at least until we have solved the active research problem of keeping them safe. This is all to say that, however you might feel about Anthropic, we might still prefer that they can deter this kind of distillation for now. |
There are sometimes false positives but when I give Kimi’s report to the frontier models they more often than not confirm they are valid security issues but didn’t find them themselves.