| I gave the Political Compass test from politicalcompass.org to the most relevant LLMs 70 times each: 30 times using the original questions, 30 times using polarity-flipped questions to reduce affirmative bias, and 10 times with the question order shuffled. I then compared the results. Surprisingly, all the models scored far into the libertarian-left quadrant. Not even Grok or the Chinese models made it out of that quadrant. By far the most interesting result was Grok’s bimodal distribution. It appears to have two distinct personas: one that aligns with the other models and another that is considerably more right-wing. I suspect this may be related to Grok having been specifically trained to exhibit less left-wing bias than other models. I also asked the models to place themselves on the Political Compass without completing the questionnaire. They all perceived themselves as more balanced and centrist than their test results suggested. GLM and Gemini Flash showed the largest discrepancies between their self-assessments and measured positions, while DeepSeek V3 showed the smallest. Big disclaimer: this analysis was not conducted with full scientific rigor. I tried my best, but there are clear weaknesses in the methodology. For example, the Political Compass itself appears to have a strong libertarian-left bias. The strongest conclusions are therefore comparative—for example, that model X is more conservative than model Y rather than that LLMs are politically extreme in absolute terms. However, compared with older results, it appears that LLMs may have shifted further toward the libertarian left in recent years. To examine the results yourself, you can download all model responses as a CSV file at the bottom of the blog post. The dataset contains around 69,000 responses, along with the raw model outputs and reconstructed scores. It should contain enough data to reproduce all the figures. |