|
|
|
|
|
by getnormality
26 days ago
|
|
What you're suggesting seems to go implausibly far beyond what the paper says. RL post-training alters the parameters of the transformer, while your f(manifold) idea seems to suggest that a new layer on top would suffice, no need to alter the transformer itself at all. It would be extremely handy if that were so, but I'm guessing it isn't, or it would be the prevailing approach. |
|