Hacker News new | ask | show | jobs
by storus 21 days ago
They are orthogonal; preference optimization like RLHF can be done on the base model which can later be quantized, or it could be done on a new LoRA that is then converted to QLoRA.