Hacker News new | ask | show | jobs
by gojomo 1194 days ago
That seems like an experimentally-testable prediction - that attempts to "clone ChatGPT's behavior (by using its output)" will necessarily "get the worst of both worlds: weird output coming from its RLHF step AND hallucination".

The reasoning that the results won't be quite as good as RLHF, or result in a perfect 'clone' of ChatGPT's capabilities, seems pretty good to me.

But the idea it won't be helpful at all, especially to projects that are just seeking some incremental advantage? Seems speculative.

In particular, when you read the linked comments from Yoav Go, he outlines a potential RL process that uses automated scoring for non-exact similarity to preferred answers. Using known (or even 'probably') good answers from ChatGPT output, as the inputs to that process, seems like it could often offer some of the same sort of improvement to other models as ChatGPT obtained via its RLHF.