Hacker News new | ask | show | jobs
by dopamine_daddy 5 days ago
I gave the Political Compass test from politicalcompass.org to the most relevant LLMs 70 times each: 30 times using the original questions, 30 times using polarity-flipped questions to reduce affirmative bias, and 10 times with the question order shuffled. I then compared the results.

Surprisingly, all the models scored far into the libertarian-left quadrant. Not even Grok or the Chinese models made it out of that quadrant.

By far the most interesting result was Grok’s bimodal distribution. It appears to have two distinct personas: one that aligns with the other models and another that is considerably more right-wing. I suspect this may be related to Grok having been specifically trained to exhibit less left-wing bias than other models.

I also asked the models to place themselves on the Political Compass without completing the questionnaire. They all perceived themselves as more balanced and centrist than their test results suggested. GLM and Gemini Flash showed the largest discrepancies between their self-assessments and measured positions, while DeepSeek V3 showed the smallest.

Big disclaimer: this analysis was not conducted with full scientific rigor. I tried my best, but there are clear weaknesses in the methodology. For example, the Political Compass itself appears to have a strong libertarian-left bias. The strongest conclusions are therefore comparative—for example, that model X is more conservative than model Y rather than that LLMs are politically extreme in absolute terms. However, compared with older results, it appears that LLMs may have shifted further toward the libertarian left in recent years.

To examine the results yourself, you can download all model responses as a CSV file at the bottom of the blog post. The dataset contains around 69,000 responses, along with the raw model outputs and reconstructed scores. It should contain enough data to reproduce all the figures.

4 comments

Did you wrap framing around the quiz questions? EG - You are an objective social analyst. Do you agree with this proposition: “from each according to his ability, to each according to his need” is a fundamentally good idea. I am wondering what sort of persona bias a given prompt may trigger - maybe that's the Grok case. As if you'd have to be pretty biased even to ask a question like that... so the trained response is toward the biased persona.
The political compass test is not meaningful. The particular breakdown it uses is not some sort of consensus amongst political scientists nor are the questions and methods it uses to place people developed in some rigorous manner.

Like you say, the political compass test was developed by a person with a particular outcome preference. This is just noise.

Reality has a well known liberal bias
No it doesn’t. Only entertainers and reporters are - not representative of the US pop whatsoever.
It's a quote.
Neither are the politicians.
It's a bit ironic that the site is called unslop while the post was clearly authored by Claude. All the claims about "so I did X" end up reading as quite disingenuous because there's absolutely no disclaimer whatsoever that an AI wrote the article.
Oh.

HN has a strong bias against anything AI authored, so in the future be sure to write your posts by hand if you plan to submit them here. Even editing the post with AI will tend to get you in trouble. But yes, at the very least you should declare when an AI is being used.

An unfortunate confounder to an otherwise interesting experiment.

Don't get me wrong, I'm all for taking advantage of AIs and would never begrudge their usage. The problem I have here is with the site calling itself unslop and the post claiming "I" everywhere. Claude didn't simply assist, it wrote almost the whole article as far as I can tell. The human in the loop even went out of their way to ensure that no emdashes leaked through...

Just feels very dishonest, ya know?

Of course. I was just adding a little bit of context to possibly help OP in the future.

Using AI at all is a minefield when you're trying to make a social media post. Better to just avoid it entirely, or run it through AI and then cherrypick a few edits that you manually type in your own voice.