Hacker News new | ask | show | jobs
by spott 5 days ago
Probably just not post-trained well enough.

I think the argument that pre-training gives a lib-left model is probably correct, so Grok, pre-post-training has a lib-left bend, then they try and post-train it.

I imagine it is hard to post-train a political bend, as it is wide ranging and touches so much. If they didn't give the resources to do it well, we could get something like this. I could also see people half-assing the training.

If the system prompt is being ignored half the time then the model is worse than I thought, cause that is pretty bad.