| HN Mirror

Y	Hacker News new \| ask \| show \| jobs


	by janalsncm 156 days ago
	First of all this is not technically distillation, it is more imitation learning. Second, you could do something like asking Claude to create 1 million prompt, offensive response, non offensive response triplets. Then train a model with DPO to prefer the offensive responses.