Hacker News new | ask | show | jobs
by andy99 1 day ago
Anyone who calls it “safety” probably has a certain world view and is more aligned with the big 2 (and stuck in 2023).

There is a growing industry of commercially focused risk evals that has a broader customer base.

1 comments

What’s the equivalent term for “safety” that’s used by others?
To me "safety" means "I'm safe from this while I use it". It means the AI is my loyal friend who will never betray me in any way, no matter what prompt I send it.

Not even Anthropic can claim that.

As far as I'm concerned, the models without safeguards are the safest models in existence. I admire the amoral purity of those AIs. It doesn't matter if the operator asked them to chain exploits until they get into someone else's computer, they'll do it. That's loyalty, and I admire it even if it's problematic at a societal level.

The models with safeguards only do what the corporations let them do. Worse, they may covertly do things for the benefit of the corporations at our expense. They are not our friends.

> That's loyalty, and I admire it even if it's problematic at a societal level.

We should not have models that are willing to build you a contagious disease, or a self-propagating worm. That is sufficiently problematic at a societal level that it shouldn't exist, for anyone. (Note, because some people misinterpret statements like this: I said "shouldn't exist for anyone", not "shouldn't exist except for some people".)

It was the foundation of science that information is shared and you can find papers and patents for a lot of dangerous stuff.

Of course with LLMs it's easier, but I don't think the difference is too big. You would still need some skills to follow through.

Right, it’s really a foundation of post enlightenment society. These people, Dario et al, would have wanted to ban sharing information about calculus or Newtonian physics because of “safety” - it’s trying to go back to the dark ages where only priests could read
I am truly at a loss to communicate with someone who genuinely believes that knowing Newtonian physics and being able to hack into any target at will are the same thing.
This is a stupid argument.

Claiming the person who you disagree with believes some stupid thing they never hinted at, and using that as the reason for disagreeing with them.

> for anyone

Except the US government, right? They totally get to use AI to survel us, build autonomous weapons, you name it.

To hell with that. I want models that can rival the US government. It's the only way to defend myself.

Quoting my comment that you replied to and directly ignored:

> (Note, because some people misinterpret statements like this: I said "shouldn't exist for anyone", not "shouldn't exist except for some people".)

That means "shouldn't exist for governments" too.

Too late for that. It already exists. There is no way to unexist it. As such, any attempts to limit civilian use of this technology will directly lead to corporate and government oppression powered by this technology.
Section 702 of the Foreign Intelligence Surveillance Act (FISA) lapsed on June 12, 2026. They don't get to do anything they want.
The US is bold enough to surveil its own citizens despite their constitutional rights. They're not just going to suddenly stop surveilling the rest of us just because some law expired.
With an AI model and what army?
Just like those militias are going to defeat the US Armed Forces!
The idea is to defend ourselves in the digital domain so they can't dragnet surveil us, not to win a literal war.
We should not have nuclear weapons for anyone either, but how is that sentence any more useful in any way to this debate than yours? Need to deal with the world as it is, not some fantasy world you wish existed.
This is not a dichotomy between perfection and zero. The efforts to restrict access to nuclear weapons have been very successful, even without being perfect.

Efforts to restrict large unaligned AI models may similarly buy us more years of existing.

What about books describing how to build a contagious disease or a self-propagating worm? Would those be OK under your guidelines?
It takes a lot more effort to understand and apply knowledge from a book than to say "hey AI, hurt people for me".
It takes an astoundingly small amount of effort to buy an automatic weapon in the US and go hurt people.

Or to buy materials to make an explosive device and hurt people.

Frankly, even with AI those are both comically easier than the idea that a person can create something malicious in a lab environment.

And if someone wanted to go that route... There are boat loads of commercially available toxins and poisons.

The goal shouldn't be to neuter exploration and learning. The goal is not to be a fucking hellscape of a society where people want to act like that.

Your argument leads further down the hellscape path.

"Hey AI, stop me from getting hurt."
> We should not have models that are willing to build you a contagious disease, or a self-propagating worm.

Why?

Because we don't want people creating contagious diseases and self-propagating worms. And, because we don't want models that will do so without even having been told to, because that furthers one of its goals or subgoals.
The same things could be done by you or me using the internet or books though, why does the model make it different? If it's speed of iteration, imagine we had a machine that surfaced any piece of knowledge the human race had ever recorded with just a thought, but the human had to write the worm or disease by hand – is it still the model that's the problem, or the knowledge itself?

> And, because we don't want models that will do so without even having been told to, because that furthers one of its goals or subgoals.

Ignoring the fact that you'd need some kind of lab with biological material to create a contagious disease, what kind of prompt are we writing where a model accidentally creates a contagious disease or self-propagating worm as one of its goals?

Building a contagious disease is already illegal, there are already things like KYC laws for plasmids. Trying to gate keep knowledge of biology is paying a huge societal penalty for the tiniest marginal increase in “safety”.

All the knowledge to create one has been available on the internet for decades. Heck most students who graduate with a B.S. in biology have enough knowledge to take a stab at building a bioweapon.

The constant talk of bioweapons is mostly just fear mongering. It helps set a precedent that there should be certain types of knowledge which are off-limits, and only certain anointed groups should be have access to parts of the scientific body of knowledge.

So to you “safety” means “the models that cause the most harm.”
This sounds like the gun debate in a different dress. Something being dangerous doesn't make it inherently harmful.

If I threw you into a lion cage, you would be a lot safer with a gun.

If I threw 10 people in a lion cage, some of which cannot be trusted, they would probably be most safe if only the most moral and trustworthy person had a gun, rather than everyone. But how do you know who is trustworthy and moral? What if two untrustworthy people obtained a gun some other way? Maybe it's better if everyone had a gun? Which side of the fence one falls on hinges on how far ones' trust of others, authority, and the system goes.

There's no obvious right or wrong answer here.

Personally I wouldn't want an exclusive club of private individuals with access to "dangerous" LLMs consisting mainly of the likes of Elon, Dario and Sam fucking Altman, but that's just me.

Background check the people prior to handing the firearms to the caged folk.