Hacker News new | ask | show | jobs
by danielvaughn 4 days ago
Armchair analysis here, of course, but I'd bet that it's due to the sources that model training tends to assume as authoritative. If you look broadly at the material produced by the (American) academic world, the corporate world, and mainstream journalism over the last 10-15 years, much of it leans generally to the left. I'd imagine it would be difficult to counteract the bias without accidentally introducing an alternate bias, but I'm not really familiar with model training so I'm not sure.
2 comments

The author cites a study (Rozado, 2024, in PLOS ONE) that supposedly shows the base models (before assistant-specific post-training) basically scored at the center (0, 0).

https://unslop.run/blog/assets/exp13/f7_context.png

Yeah, and who picks the sources?

Most of the data on the internet is UGC, and most of the time it's VERY right-leaning. It was one of the reasons OpenAI was "afraid" to release GPT-2: because it produced results that people in Silicon Valley didn't like.

I remember when Microsoft released the first proto-LLM and users always steered to being a nazi.