Hacker News new | ask | show | jobs
by cgorlla 2 days ago
We discuss this in the writeup. While we expected this result, it is important for there to be data backing the claims, and an experimental setup that mirrors productions tasks is a useful tool for the conversations going on about this.
1 comments

Yeah, the benefit of showing this seems obvious to me. I probably would've expected the censorship to transfer slightly given the anthropic owl paper from years ago https://alignment.anthropic.com/2025/subliminal-learning/

But that was about transferring from a finetuned model to another finetune of the same base model, good to see more evidence that it doesn't transfer cleanly across different base models in a more realistic scenario than an "owl-loving model"

Edit: From *year ago. It's been a long year haha