Hacker News new | ask | show | jobs
by mlinsey 18 days ago
I have some sympathy to geohot's view when it comes to pure informational chatbots. It's a first amendment issue, I'm allowed to write and read books that are useful to getting away with crimes, etc.

This obviously doesn't work at all when the agents start doing real things in the real world, though. "Hey AI, I don't like my neighbor, find an exploit in the firmware for his car and make the cruise control malfunction and crash him next time he gets on the highway". This is committing a crime, not just talking about theoretical crimes. The AI can and should refuse it.

I think he's anticipating and discarding this objection with his introduction, which otherwise feels disconnected from the rest of the article. FWIW, I have changed a bike tire and I'm pretty sure most of the MTS at the big labs could. This sort of "they're just bookworms who don't understand the physical world" rhetoric aside, we are currently seeing a ton of effort and expense go towards giving the AI agents hooks into being able to perform as many real-world-consequential actions as possible. And you can do a surprising amount with just bits, from writing code to breaking into systems to sending some combination of emails, phone calls, and currency to instruct meatspace humans to do things, etc.

4 comments

> The AI can and should refuse it.

This leads to LLMs refusing to do security work for anyone, because it can’t tell if it’s being done for good or evil purposes.

Which is precisely what we’re going through right now with frontier models and it’s terrible.

These proposals always assume some perfect mechanism for identifying the thought crime with triggering on the normal requests. The people who want to commit crimes just persist until they jailbreak the guardrails while the rest of us suffer with denials.

Also if you can’t imagine these guardrails being used by governments to control inconvenient speech, you probably need to think a little harder about the realities of how these will be used.

> Which is precisely what we’re going through right now with frontier models and it’s terrible.

terrible is a bit far, but pretty hilarious for sure

> This leads to LLMs refusing to do security work for anyone, because it can’t tell if it’s being done for good or evil purposes.

Oh no, of only we had some sort of ideas about automated authentication and authorisation! (If big AI actually make a model that would reliably reject dangerous queries in the first place, they can just use auth tech to serve the current ones to vetted customers)

> If big AI actually make a model that would reliably reject dangerous queries in the first place...

Not happening, for two reasons

1) LLMs mix attacker-controlled data into their instruction stream, and it doesn't look at all like the big LLM providers consider making that impossible to be an actual priority.

2) Intent is what distinguishes "dangerous" queries from prosocial ones. Reliably determining intent is so complex that not even humans can do it.

> they can just use auth tech to serve the current ones to vetted customers

Your solution is to require everyone to surrender their privacy and then allow some company or government decide who is worthy of being allowed to use AI?

It’s crazy how authoritarian, controlling, and anti-privacy the anti-LLM conversation is getting.

I don’t think it’s a privacy issue. All the best auditors collaborate to agree on ethical auditing standards requiring complete disclosure; that’s not a privacy violation, even if it means some businesses end up effectively forced to reveal details they’d prefer to keep private.
Not sure how "some users get special access after special scrutiny" means "all users must submit to scrutiny", obviously you could just perform the scrutiny upon request.
> "Hey AI, I don't like my neighbor, find an exploit in the firmware for his car and make the cruise control malfunction and crash him next time he gets on the highway". This is committing a crime, not just talking about theoretical crimes. The AI can and should refuse it.

The obvious problem being that the user doesn't have any need to provide the context that they're trying to commit a crime and can just ask how to do something without providing a reason or making one up.

At which point you'd have the model trying to impute a reason and often getting it dangerously wrong, e.g. refusing to disclose a vulnerability when the user is actually the defender who needs to patch/mitigate it.

"The AI" cannot "refuse" any more than a hammer can refuse. If you use a hammer to kill someone, you commit a crime and the hammer is unprosecuted. If you design a hammer to kill people and then give it away, you are partially liable for the deaths it causes. If you make a normal hammer and someone uses it to kill, you bear no blame.

I don't see why these standards should change when the hammer also emits text messages.

> If you design a hammer to kill people and then give it away, you are partially liable for the deaths it causes.

This doesn't seem to be the case today (weapons manufacture).

Big AI aren't selling hammers though, they are renting them out. And if you rent a hammer to someone asking how hard to you need to hit someone over the head to kill, you are likely to be found liable.
The AI can absolutely refuse, what are you talking about?

These things can and do say “sorry I won’t help you with that” based on the nature of your request. What do you call that if not a refusal?

> It's a first amendment issue, I'm allowed to write and read books that are useful to getting away with crimes, etc.

I'm neither American nor a lawyer.

Is "conspiracy" protected under the first amendment?

If you discuss a crime with someone to learn about it, does that count as "conspiracy"?

GP is correct, books describing, advocating, even instructing crime in the abstract are almost surely protected speech.
> If you discuss a crime with someone to learn about it, does that count as "conspiracy"?

Replace 'conspiracy' with 'agreement'. Discussing a crime is not a conspiracy. Agreeing with someone to commit the crime is.

Yes conspiracy is covered, now if you act it out… that’s a different story.