Y
Hacker News
new
|
ask
|
show
|
jobs
by
storus
1 day ago
What would be the current best method to fine-tune it for my own specific agentic tasks? LoRA + DPO? GRPO? Something else?
1 comments
whimsicalism
1 day ago
LoRA + SFT, but it'll be big - better to wait for a finetuning API from one of the providers, I wouldn't jump straight to RL or off-policy pseudo-RL like DPO.
link