Hacker News new | ask | show | jobs
by SpicyLemonZest 18 days ago
The point is that we need to have a safety plan in place before an AI is smart enough to radically reshape the world. If you've got an AI that's ready to start sending humanity into the next era of civilization, it may be too late to control when and how it does that.

> Heck, a proof for P=NP or P!=NP or solve the The Riemann Hypothesis. Just give me something truly exciting and I will believe AGI is around the corner, until then I will see it as cool technology, that while beneficial to me, also helped cause the biggest amount of disinformation we've every seen.

I hope you'll keep this in mind when those milestones are reached. What I've seen a lot of people do, unfortunately, is pretend that the impressive things nobody thought AI could do 5 years ago are trivial things that aren't very hard.

3 comments

Yes, but why do we think the likelihood of any of that happening is great enough to warrant all this effort and cost
Because people who predicted the AI capabilities we've seen get developed over the past decade also predicted that dangerous AI systems capable of these things would follow soon after.

It's not a settled debate, even among experts, and perhaps in retrospect we'll realize AI safety was unnecessary or based on fundamental confusions. But if the median ergonomics researcher gave a 5% chance that a new chair will be so comfortable that it drives humanity extinct (https://www.nature.com/articles/d41586-024-00147-z), I would definitely want the government to start measuring and regulating chair comfort, even if that was costly and even if that meant I couldn't buy a comfy new chair I wanted.

Experts as in AI-2027, Jenson “I think we’ve achieved AGI in March 2026” Huang, or that AI Researcher(tm) calling to pre-emptively bomb Microsoft data centers in summer 2022?
A safety plan. It doesn't have to be "a few very smart people detached from reality convincing themselves they're messiahs that must keep the tech from the Bad Guys(tm), unwashed masses, and a runaway, because they think their interpretation of their own sci-fi lore is the only possible course of events"

>I hope you'll keep this in mind when those milestones are reached.

The problem is that people in charge of AI keep making self-fulfilling prophecies. Just like with any research, if you want to find something sensible in the cloud patterns, you will.

What is an example you have in mind of a self-fulfilling prophecy? I genuinely don't know what you could be referring to. It seems to me that they keep making surprising prophecies, and the popular reaction to them seamlessly transitions from "that's crazy, no way it will happen" to "that's silly, it's just a cloud pattern". Did you find it obvious or self-fulfilling in 2025 that LLMs would soon be able to resolve open questions in mathematical research?
Self-fulfilling prophecies are social effects, not real predictions. Anthropic's "We predict our model misbehavior potentially being able do destroy the world in 1% of the cases" -> "We want to find the evidence of our model misbehaving, and we want it bad" -> "See, our model is hacking our rewards and has functional emotions, this means it's misbehaving with the intention of destroying the... HUMANITY!1" -> repeat x100, manipulate the media into amplifying it x10000 for clicks -> people are begging to safeguard them from the evil AI. Which is already likely to happen, the average layman's Overton window already includes the fantasy of rogue AI.

None of that was real or remotely dangerous in the first place, of course. It wouldn't have resulted in controls, had they not been scaremongering. This will end in extreme fascism or people getting enslaved "for their own safety", and it won't even require malicious intent, only incentives, detachment from reality, and confirmation bias. Although it doesn't exclude malice either.

Why would we be able to design a safety plan that would control such a powerful force? It would be like putting umbrellas on an asteroid hoping to slow its fall. That sounds like delusions of grandeur.
It'll certainly require a bit more than umbrellas! NASA has spent hundreds of millions of dollars developing asteroid defense systems, including a proof of concept for redirection in 2022. So, you know, let's try at least that much before declaring there's probably no way to do it.