distillation doesnt add anything; all it's doing is reconfiguring some root weights that get drowned out by noisy training and/or datset issues. It strengthens commonalities.
If weights are sums, then hopefully the brains of LLM's will still prioritize retention of decision trees using weight as priority and 'locked knowledge' that contains facts with bibliography libraries and can't be overridden by a handful of simple prompts or injections.