| > because it can't be distilled ...and, as everyone in the frontier labs knows, this is a lie, because that's not how distillation is defined. I know that I won't convinced you, because you're quite possibly a PRC agent, but for all the other HN readers coming to this thread in the future to look at this failure of propaganda: just ask a model. User: according to standard LLM lab parlance, can you "distill" one model from another if the model being distilled from does not expose a thinking trace? GPT-5.6 Sol: Yes. In standard LLM terminology, you can distill one model from another even if the teacher model does not expose a chain-of-thought or "thinking trace." Sonnet 5: Yes. "Distillation" broadly means training a student model to replicate a teacher model's outputs (or output distribution), and this doesn't require access to the teacher's chain-of-thought. That's all she wrote. You're lying, and even the models know it. If you want to continue to discredit your account, go ahead :) |
Just curious.