Hacker News new | ask | show | jobs
by qeternity 12 days ago
It's a defensive tactic to reduce the effectiveness of distillation.

Say of that what you will, but it's not because they want to wrest control from users.

It's because they don't want Chinese companies to do exactly what Moonshot (Kimi creators) and others have done.

3 comments

Anthropic’s position being that it is entitled to train models on the creative works of anyone at any time, but its own slop generators’ outputs are sacred jewels that must be protected from being learned from.
Its not surprising a thief would be the most paranoid in securing their spoils.
According to Pat Toulme, the thinking traces and outputs that the Chinese researchers distilled are useful for getting initial trajectories, to prevent a cold start during RL. Once you get those initial correct trajectories (the model actually solving a task), you can generate you own new traces, so the distillation is already done, there is no need for further reasoning traces. Sure they could probably get better alignment with the frontier by distilling fruther, but in any case the damage is already done, at this point hiding reasoning is mostly just hurting users
It's funny because Anthropic is harming my use of their product, to not even stop the supposed theft of their thought chains. Something that any other supposed upstarts could presumably grab from Moonshot or whatever now.

And my complaint also extends to all their tool use explanations, or rather the lack thereof. I get prompted continually for tool use that I can't examine, that's a poorly formatted 1kb bash script, etc. The PM desire to hide valuable information while requiring extensive interaction has really driven the product into a very unusable place compared to where it was a few months ago. (Or perhaps that's just Claude responding to its memory of my use, and I have somehow driven it to be excessively verbose and difficult to use, which would be unfortunate... Perhaps there's a way to reset the memory.)