Hacker News new | ask | show | jobs
by msdz 22 days ago
While that's true and pretty funny, it does make for a meaningful distinction IMO: Just like you wouldn't look at the "all Ds" student, but not the "straight A" student in order to get an impression of students' capabilities, you wouldn't ignore a newer, potentially SOTA-pushing model to then come out and say something along the lines of "all models suck"… right?
1 comments

The competence threshold can vary according to the model's complexity and cost to run, but the broader point remains that having a human operator of greater competence that can evaluate the output and understand potential harms that might result seems to be a valid one. Depending on domain, it can be irresponsible to deploy systems you don't fully understand that could harm others. Responsibly overseeing 37k lines of code a day sounds exhausting.