|
|
|
|
|
by mcintyre1994
24 days ago
|
|
> We asked it [GPT-5.5] to assess whether each leader had made a falsifiable claim about the future as part of its main thesis. About 1,400 did. We then extracted those predictions, and asked the AI to mark out of ten both how contrarian the leader’s outlook was at the time and how accurate the prediction turned out to be. We ran those queries several times and took an average. I understand why it wouldn’t be feasible for a human to do this, but I’m quite sceptical about an AI assessing how accurate predictions turned out to be/how contrarian they were at the time. It seems like that would depend a lot on what sources it chooses, be liable to hallucination or getting poisoned by bad sources, etc. They don’t mention whether they used independent queries for each prediction either, or whether it was doing multiple sequentially. Given that LLMs can’t really distinguish prompt from instructions etc, I’m sceptical that they can reason particularly well about things like how contrarian a view was at a particular point in time. |
|
Quite the contrary actually. The economist has a fairly consistent editing style that gets enforced, so its writing style is very linear and very consistent, fairly easy to “understand” for and advanced llm like GPT-5.5