Hacker News new | ask | show | jobs
by operation_moose 22 days ago
AI feels to me like having access to someone who got a D in literally every single course offered at a university. If you don't know anything about the subject they are smarter than you. If you do know the subject its unsettling how bad they are. Basically the Gell-Mann effect:

> The phenomenon of a person trusting newspapers for topics which that person is not knowledgeable about, despite recognizing the newspaper as being extremely inaccurate on certain topics which that person is knowledgeable about.

They've improved from someone who failed every single university course a couple years ago. Maybe they'll get to a C or even a B in the future; maybe not.

4 comments

It's more like giving an army of crack heads with a heap of motivation access to textbooks and the internet to do your bidding
>AI feels to me like having access to someone who got a D in literally every single course offered at a university. If you don't know anything about the subject they are smarter than you. If you do know the subject, its unsettling how bad they are. Basically the Gell-Mann effect:

I don't know how you can say that with a straight face when AI's capable of matching the best human students on the hardest exams we have, like the IMO, the Putnam and the bar exam. I can only assume you've only ever used the free tier of any AI service.

Exams are usually discussed and part of the corpus though right? Performance on novel problems are really the only good metric and those decay really quickly across models once they’re known, like the pelican on a bicycle test
I like that you start by saying that you don’t know how someone could draw that conclusion about AI and end by saying they probably got there by using AI
While that's true and pretty funny, it does make for a meaningful distinction IMO: Just like you wouldn't look at the "all Ds" student, but not the "straight A" student in order to get an impression of students' capabilities, you wouldn't ignore a newer, potentially SOTA-pushing model to then come out and say something along the lines of "all models suck"… right?
The competence threshold can vary according to the model's complexity and cost to run, but the broader point remains that having a human operator of greater competence that can evaluate the output and understand potential harms that might result seems to be a valid one. Depending on domain, it can be irresponsible to deploy systems you don't fully understand that could harm others. Responsibly overseeing 37k lines of code a day sounds exhausting.
We're trying to up the GPA on the code quality exam by giving them access to structural metrics calculated on graph representations of the programs they write. Hoping this is the study guide template they need to start getting an A in minimizing slop, even if they're actually a D student. They're good at optimizing against scores...

https://github.com/Krv-Labs/topos

Good on you for trying, and I'm not saying it can't improve things. Optimizing for structural metrics is still only as good as the metrics. I can say with confidence that no metric exists for good design, because good design is art. Art requires creativity and the evaluation lies in the eye of the beholder. It also doesn't help with the bigger picture, AIs often do the wrong thing because they have been tasked to do the wrong thing.
There are no studies to suggest that "Gell-Mann" is a real phenomenon. It was invented by an author, Michael Crichton, on a whim.

Any rational reader will adjust their faith in a publication based on the identified errors it makes. Too many and they'll reject it as a source for anything beyond "a thing may have happened".

I thought you generally knew what you were talking about till i read this. Now I am wondering if you know anything at all.