Hacker News new | ask | show | jobs
by logicchains 22 days ago
>AI feels to me like having access to someone who got a D in literally every single course offered at a university. If you don't know anything about the subject they are smarter than you. If you do know the subject, its unsettling how bad they are. Basically the Gell-Mann effect:

I don't know how you can say that with a straight face when AI's capable of matching the best human students on the hardest exams we have, like the IMO, the Putnam and the bar exam. I can only assume you've only ever used the free tier of any AI service.

2 comments

Exams are usually discussed and part of the corpus though right? Performance on novel problems are really the only good metric and those decay really quickly across models once they’re known, like the pelican on a bicycle test
I like that you start by saying that you don’t know how someone could draw that conclusion about AI and end by saying they probably got there by using AI
While that's true and pretty funny, it does make for a meaningful distinction IMO: Just like you wouldn't look at the "all Ds" student, but not the "straight A" student in order to get an impression of students' capabilities, you wouldn't ignore a newer, potentially SOTA-pushing model to then come out and say something along the lines of "all models suck"… right?
The competence threshold can vary according to the model's complexity and cost to run, but the broader point remains that having a human operator of greater competence that can evaluate the output and understand potential harms that might result seems to be a valid one. Depending on domain, it can be irresponsible to deploy systems you don't fully understand that could harm others. Responsibly overseeing 37k lines of code a day sounds exhausting.