| HN Mirror

Y	Hacker News new \| ask \| show \| jobs

by smusamashah 555 days ago

These kind of tests (or may be all tests) should show *success rate* instead of a single pass/fail.

I believe Claude or even Gemini can succeed if system prompt is improved e.g. tell it to re-evaluate it's answer before finalising, can even tell it to do "thinking" within <thinking> tags. I use claude like that and it often goes over it's answer and corrects itself within same reply. On the other hand it can also incorrectly assume it made a mistake and can sometimes uncorrect itself.

Edit: Using o1's step by step problem solving example from OpenAI blog post made Claude go step by step in similar depth too. Could even do that here to get better success rate in non-o1 models.