| HN Mirror

Y	Hacker News new \| ask \| show \| jobs

by teruakohatu 1174 days ago

Interesting.

> The response: {"confidence": "very low", "en": "I'm not sure, but I don't think the moon is made of cheese."}

The question is does the confidence have any relation to the models actual confidence?

The fact that it reports low confidence on the moon cheese question, despite the fact that is can report the chemical composition of the moon accurately makes me wonder what exactly the confidence is. Seems more like sentiment analysis on its own answer.

2 comments

lsy 1174 days ago

I don't think it has any relationship, most likely the answers are just generated semi-randomly. Even the one it's "very" confident about is not agreed-upon (Wikipedia says the outcome was "inconclusive"). Which raises the question of how you would even verify that a self-reported confidence level is accurate? Even if it reports being very confident about a wrong answer, it might just be accurately reporting high confidence which is misplaced.

layer8 1174 days ago

My view is that ChatGPT isn’t a singular “it”. Its output is a random sampling from a range of possible “its”, the only (soft) constraint being the contents of the current conversation.

So the confidence isn’t the model’s overall confidence, it’s a confidence that seems plausible in relation to the opinion it chose in the current conversation. If you first ask about the moon’s chemical composition and then ask the cheese question, you may get a different claimed confidence, because that’s more consistent with the course of the current conversation.

Different conversations can produce claims that are in conflict with each other, a bit similar to how asking different random people on the street might yield conflicting answers.