Hacker News new | ask | show | jobs
by cgorlla 2 days ago
The examples you're talking about are not involved in the training process, so their number is irrelevant. As stated in the post, the goal of this work is to determine whether a teacher's unrelated behaviors are inherited by the student distilled on a different task. Changing how the model thinks about the Holodomor is completely irrelevant.
1 comments

> Changing how the model thinks about the Holodomor is completely irrelevant.

Your post title is literally "Distilling DeepSeek into GPT-OSS doesn't transfer censorship."

Like I'm not really interested in debating you on this because even the title is nonsense, there is no good faith interpretation of what you're doing here.

Distillation is such a wide concept, and you have such a narrow domain, it's not an even somewhat useful experiment to make the claim that you're making.

The fact that you literally thought the examples were used in SFT in your last comment ago calls into question the utility of this conversation, notwithstanding the implication that those examples were used to improve…financial performance?

This is a very standard setup for a distillation problem. The vast majority of companies don't care about the "wide concept", this is what most distillation consists of. They want to improve models on a narrow domain. It should be understood that this is by and large a low risk vector for this sort of behavior to transfer. That is what we are measuring, and we are very open about it.

> Testing whether censorship transmits through unrelated data requires that it never appear in the data.

> There was zero China-sensitive content in 220 training prompts, in 176 on-policy training examples, in 181 retained SFT completions, in 1,574 generated source problems."

It's really not my fault you wrote a rambling article and while skimming (best an article earns out of me with an off-smelling title) I took that to imply there was an SFT step in your distillation pipeline.

Maybe the AI that wrote the article for you was a bit confused on that as well?

-

Also I question your understanding of this thread if you're wasting so many words trying to explain distillation to me.

(I mean, you're wrong btw. If we're going full pedant then most compute spent on distillation is labs very broadly distilling their own models into smaller models that are still pretty damn large and expensive to distill...)

But sure, small scale distillation is usually for narrow domain specific tasks, welcome to 2019. The entire point of this thread is that "distillation" for such narrow use cases couldn't reasonably affect censorship without intention.

You can introduce misalignment even with a very narrow focus (https://arxiv.org/html/2502.17424v2), but it doesn't happen by accident.

So if your goals with distillation weren't centered around censorship, and weren't meant to introduce censorship, then why are you trying to draw this tenuous link?

I guess the AI that wrote your comment for you also conflated the SFT step of the target domain with the political prompts, which, in the sentence you quoted, contradicts your original comment...